{
  "id": 283916,
  "title": "[tl;dr summary list] Research Papers summary about Text & Image matching",
  "url": "/competitions/wikipedia-image-caption/discussion/283916",
  "author_name": "",
  "post_date": "2021-10-28T14:18:28.213062800Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Below is a list of relevant papers about text &amp; image matching that I went skimmed through while working on this competition. </p>\n<p><strong>Treating \"baseline\" as:</strong></p>\n<ul>\n<li>simple aggregation of the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions.</li>\n</ul>\n<hr>\n<p><strong>Graph Structured Network for Image-Text Matching</strong></p>\n<ul>\n<li>present a novel Graph Structured Matching Network (GSMN) to learn fine-grained correspondence. </li>\n<li>learns object, relation and attribute as a structured phrase - allows to learn correspondence of object, relation and attribute separately. </li>\n<li>training by node-level matching and structure-level matching.<br>\n<strong>Code:</strong> <a href=\"https://github.com/CrossmodalGroup/GSMN\" target=\"_blank\">https://github.com/CrossmodalGroup/GSMN</a>.<br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/pdf/2004.00277.pdf\" target=\"_blank\">https://arxiv.org/pdf/2004.00277.pdf</a></li>\n</ul>\n<hr>\n<p><strong>Stacked Cross Attention for Image-Text Matching</strong></p>\n<ul>\n<li>present Stacked Cross Attention to discover the full latent alignments using both image regions and words in a sentence as context and infer image-text similarity. </li>\n<li>Achieves the state-of-the-art results on the MSCOCO and Flickr30K datasets. On Flickr30K<br>\n<strong>Code:</strong> <a href=\"https://github.com/kuanghuei/SCAN\" target=\"_blank\">https://github.com/kuanghuei/SCAN</a><br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/abs/1803.08024\" target=\"_blank\">https://arxiv.org/abs/1803.08024</a></li>\n</ul>\n<hr>\n<p><strong>Similarity Reasoning and Filtration for Image-Text Matching</strong></p>\n<ul>\n<li>propose a novel Similarity Graph Reasoning and Attention Filtration (SGRAF) network for image-text matching. </li>\n<li>first: vector-based similarity representations are learned. then: Similarity Graph Reasoning (SGR) using graph convolution infers relationaware similarities.</li>\n<li>also used: similarity attention filtration (SAF) to integrate these alignments<br>\n<strong>Code:</strong> <a href=\"https://github.com/Paranioar/SGRAF\" target=\"_blank\">https://github.com/Paranioar/SGRAF</a><br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2101.01368\" target=\"_blank\">https://arxiv.org/abs/2101.01368</a></li>\n</ul>\n<hr>\n<p><strong>Text Matching as Image Recognition</strong></p>\n<ul>\n<li>Firstly, a matching matrix whose entries represent the similarities between words is constructed and viewed as an image. </li>\n<li>Then a convolutional neural network is utilized to capture rich matching patterns in a layer-by-layer way.<br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/pdf/1602.06359.pdf\" target=\"_blank\">https://arxiv.org/pdf/1602.06359.pdf</a></li>\n</ul>\n<hr>\n<p><strong>Consensus-Aware Visual-Semantic Embedding for Image-Text Matching</strong></p>\n<ul>\n<li>propose a Consensus-aware Visual-Semantic Embedding (CVSE) model to incorporate the commonsense knowledge shared between both modalities, into image-text matching. </li>\n<li>the information is exploited by computing the statistical co-occurrence correlations between the semantic concepts from the image captioning corpus and deploying the constructed concept correlation graph to yield the consensus-aware concept (CAC) representations.</li>\n<li>then trains on the associations and alignments between image and text based on the exploited consensus as well as the instance-level representations for both modalities. <br>\n<strong>Code:</strong> <a href=\"https://github.com/BruceW91/CVSE\" target=\"_blank\">https://github.com/BruceW91/CVSE</a>.<br>\n<strong>Paper:</strong> <a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123690018.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123690018.pdf</a></li>\n</ul>\n<hr>\n<p>Best of luch guys!<br>\nCheers! ^^ </p>",
  "messages": [
    {
      "id": "1563720",
      "postDate": "10/28/2021 14:18:28",
      "content": "<p>Below is a list of relevant papers about text &amp; image matching that I went skimmed through while working on this competition. </p>\n<p><strong>Treating \"baseline\" as:</strong></p>\n<ul>\n<li>simple aggregation of the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions.</li>\n</ul>\n<hr>\n<p><strong>Graph Structured Network for Image-Text Matching</strong></p>\n<ul>\n<li>present a novel Graph Structured Matching Network (GSMN) to learn fine-grained correspondence. </li>\n<li>learns object, relation and attribute as a structured phrase - allows to learn correspondence of object, relation and attribute separately. </li>\n<li>training by node-level matching and structure-level matching.<br>\n<strong>Code:</strong> <a href=\"https://github.com/CrossmodalGroup/GSMN\" target=\"_blank\">https://github.com/CrossmodalGroup/GSMN</a>.<br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/pdf/2004.00277.pdf\" target=\"_blank\">https://arxiv.org/pdf/2004.00277.pdf</a></li>\n</ul>\n<hr>\n<p><strong>Stacked Cross Attention for Image-Text Matching</strong></p>\n<ul>\n<li>present Stacked Cross Attention to discover the full latent alignments using both image regions and words in a sentence as context and infer image-text similarity. </li>\n<li>Achieves the state-of-the-art results on the MSCOCO and Flickr30K datasets. On Flickr30K<br>\n<strong>Code:</strong> <a href=\"https://github.com/kuanghuei/SCAN\" target=\"_blank\">https://github.com/kuanghuei/SCAN</a><br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/abs/1803.08024\" target=\"_blank\">https://arxiv.org/abs/1803.08024</a></li>\n</ul>\n<hr>\n<p><strong>Similarity Reasoning and Filtration for Image-Text Matching</strong></p>\n<ul>\n<li>propose a novel Similarity Graph Reasoning and Attention Filtration (SGRAF) network for image-text matching. </li>\n<li>first: vector-based similarity representations are learned. then: Similarity Graph Reasoning (SGR) using graph convolution infers relationaware similarities.</li>\n<li>also used: similarity attention filtration (SAF) to integrate these alignments<br>\n<strong>Code:</strong> <a href=\"https://github.com/Paranioar/SGRAF\" target=\"_blank\">https://github.com/Paranioar/SGRAF</a><br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2101.01368\" target=\"_blank\">https://arxiv.org/abs/2101.01368</a></li>\n</ul>\n<hr>\n<p><strong>Text Matching as Image Recognition</strong></p>\n<ul>\n<li>Firstly, a matching matrix whose entries represent the similarities between words is constructed and viewed as an image. </li>\n<li>Then a convolutional neural network is utilized to capture rich matching patterns in a layer-by-layer way.<br>\n<strong>Paper:</strong> <a href=\"https://arxiv.org/pdf/1602.06359.pdf\" target=\"_blank\">https://arxiv.org/pdf/1602.06359.pdf</a></li>\n</ul>\n<hr>\n<p><strong>Consensus-Aware Visual-Semantic Embedding for Image-Text Matching</strong></p>\n<ul>\n<li>propose a Consensus-aware Visual-Semantic Embedding (CVSE) model to incorporate the commonsense knowledge shared between both modalities, into image-text matching. </li>\n<li>the information is exploited by computing the statistical co-occurrence correlations between the semantic concepts from the image captioning corpus and deploying the constructed concept correlation graph to yield the consensus-aware concept (CAC) representations.</li>\n<li>then trains on the associations and alignments between image and text based on the exploited consensus as well as the instance-level representations for both modalities. <br>\n<strong>Code:</strong> <a href=\"https://github.com/BruceW91/CVSE\" target=\"_blank\">https://github.com/BruceW91/CVSE</a>.<br>\n<strong>Paper:</strong> <a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123690018.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123690018.pdf</a></li>\n</ul>\n<hr>\n<p>Best of luch guys!<br>\nCheers! ^^ </p>",
      "rawMarkdown": "Below is a list of relevant papers about text & image matching that I went skimmed through while working on this competition. \n\n\n**Treating \"baseline\" as:**\n- simple aggregation of the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions.\n\n____\n\n**Graph Structured Network for Image-Text Matching**\n- present a novel Graph Structured Matching Network (GSMN) to learn fine-grained correspondence. \n- learns object, relation and attribute as a structured phrase - allows to learn correspondence of object, relation and attribute separately. \n- training by node-level matching and structure-level matching.\n**Code:** https://github.com/CrossmodalGroup/GSMN.\n**Paper:** https://arxiv.org/pdf/2004.00277.pdf\n\n____\n\n**Stacked Cross Attention for Image-Text Matching**\n- present Stacked Cross Attention to discover the full latent alignments using both image regions and words in a sentence as context and infer image-text similarity. \n- Achieves the state-of-the-art results on the MSCOCO and Flickr30K datasets. On Flickr30K\n**Code:** https://github.com/kuanghuei/SCAN\n**Paper:** https://arxiv.org/abs/1803.08024\n\n____\n\n**Similarity Reasoning and Filtration for Image-Text Matching**\n- propose a novel Similarity Graph Reasoning and Attention Filtration (SGRAF) network for image-text matching. \n- first: vector-based similarity representations are learned. then: Similarity Graph Reasoning (SGR) using graph convolution infers relationaware similarities.\n- also used: similarity attention filtration (SAF) to integrate these alignments\n**Code:** https://github.com/Paranioar/SGRAF\n**Paper:** https://arxiv.org/abs/2101.01368\n\n____\n\n**Text Matching as Image Recognition**\n- Firstly, a matching matrix whose entries represent the similarities between words is constructed and viewed as an image. \n- Then a convolutional neural network is utilized to capture rich matching patterns in a layer-by-layer way.\n**Paper:** https://arxiv.org/pdf/1602.06359.pdf\n\n____\n\n**Consensus-Aware Visual-Semantic Embedding for Image-Text Matching**\n- propose a Consensus-aware Visual-Semantic Embedding (CVSE) model to incorporate the commonsense knowledge shared between both modalities, into image-text matching. \n- the information is exploited by computing the statistical co-occurrence correlations between the semantic concepts from the image captioning corpus and deploying the constructed concept correlation graph to yield the consensus-aware concept (CAC) representations.\n- then trains on the associations and alignments between image and text based on the exploited consensus as well as the instance-level representations for both modalities. \n**Code:** https://github.com/BruceW91/CVSE.\n**Paper:** https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123690018.pdf\n\n____\n\nBest of luch guys!\nCheers! ^^",
      "votes": null
    },
    {
      "id": "1563758",
      "postDate": "10/28/2021 15:01:53",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1563758,
      "author_name": "gilangarisptr",
      "author_url": "",
      "post_date": "10/28/2021 15:01:53",
      "content": "<p>thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1563720": "Below is a list of relevant papers about text & image matching that I went skimmed through while working on this competition. \n\n\n**Treating \"baseline\" as:**\n- simple aggregation of the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions.\n\n____\n\n**Graph Structured Network for Image-Text Matching**\n- present a novel Graph Structured Matching Network (GSMN) to learn fine-grained correspondence. \n- learns object, relation and attribute as a structured phrase - allows to learn correspondence of object, relation and attribute separately. \n- training by node-level matching and structure-level matching.\n**Code:** https://github.com/CrossmodalGroup/GSMN.\n**Paper:** https://arxiv.org/pdf/2004.00277.pdf\n\n____\n\n**Stacked Cross Attention for Image-Text Matching**\n- present Stacked Cross Attention to discover the full latent alignments using both image regions and words in a sentence as context and infer image-text similarity. \n- Achieves the state-of-the-art results on the MSCOCO and Flickr30K datasets. On Flickr30K\n**Code:** https://github.com/kuanghuei/SCAN\n**Paper:** https://arxiv.org/abs/1803.08024\n\n____\n\n**Similarity Reasoning and Filtration for Image-Text Matching**\n- propose a novel Similarity Graph Reasoning and Attention Filtration (SGRAF) network for image-text matching. \n- first: vector-based similarity representations are learned. then: Similarity Graph Reasoning (SGR) using graph convolution infers relationaware similarities.\n- also used: similarity attention filtration (SAF) to integrate these alignments\n**Code:** https://github.com/Paranioar/SGRAF\n**Paper:** https://arxiv.org/abs/2101.01368\n\n____\n\n**Text Matching as Image Recognition**\n- Firstly, a matching matrix whose entries represent the similarities between words is constructed and viewed as an image. \n- Then a convolutional neural network is utilized to capture rich matching patterns in a layer-by-layer way.\n**Paper:** https://arxiv.org/pdf/1602.06359.pdf\n\n____\n\n**Consensus-Aware Visual-Semantic Embedding for Image-Text Matching**\n- propose a Consensus-aware Visual-Semantic Embedding (CVSE) model to incorporate the commonsense knowledge shared between both modalities, into image-text matching. \n- the information is exploited by computing the statistical co-occurrence correlations between the semantic concepts from the image captioning corpus and deploying the constructed concept correlation graph to yield the consensus-aware concept (CAC) representations.\n- then trains on the associations and alignments between image and text based on the exploited consensus as well as the instance-level representations for both modalities. \n**Code:** https://github.com/BruceW91/CVSE.\n**Paper:** https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123690018.pdf\n\n____\n\nBest of luch guys!\nCheers! ^^",
    "1563758": "thanks for sharing"
  },
  "source": "meta"
}