{
  "id": 272023,
  "title": "Welcome to the Wikipedia Image/Caption Matching Competition!",
  "url": "/competitions/wikipedia-image-caption/discussion/272023",
  "author_name": "Miriam Redi",
  "post_date": "2021-09-13T17:31:27.907000",
  "votes": 16,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Welcome to the Wikipedia Image/Caption competition!</p>\n<p>We are beyond excited to support this competition, and call for brilliant people like you to solve challenging problems in the vision and language space! In this challenge, you will  play with large amounts of multilingual and multimodal data, coming from the world's largest online encyclopedia. </p>\n<p>Your contributions to this competition are contributions to free and open knowledge!</p>\n<p>Thanks for joining, we are here to help should you have any questions!</p>\n<p>Miriam</p>",
  "messages": [
    {
      "id": 1511796,
      "postDate": "2021-09-13T17:31:27.907Z",
      "content": "<p>Welcome to the Wikipedia Image/Caption competition!</p>\n<p>We are beyond excited to support this competition, and call for brilliant people like you to solve challenging problems in the vision and language space! In this challenge, you will  play with large amounts of multilingual and multimodal data, coming from the world's largest online encyclopedia. </p>\n<p>Your contributions to this competition are contributions to free and open knowledge!</p>\n<p>Thanks for joining, we are here to help should you have any questions!</p>\n<p>Miriam</p>",
      "rawMarkdown": "Welcome to the Wikipedia Image/Caption competition!\n\nWe are beyond excited to support this competition, and call for brilliant people like you to solve challenging problems in the vision and language space! In this challenge, you will  play with large amounts of multilingual and multimodal data, coming from the world's largest online encyclopedia. \n\nYour contributions to this competition are contributions to free and open knowledge!\n\nThanks for joining, we are here to help should you have any questions!\n\nMiriam\n",
      "votes": 16
    },
    {
      "id": 1511803,
      "postDate": "2021-09-13T17:45:28.113Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>\n<p>Thanks for preparing this competition and releasing it on Kaggle. The problem looks really interesting but it lacks incentive to join. A relatively large dataset that requires lots of computational power and no medals or prizes can be a red flag for lots of kagglers. This competition should at least award medals.</p>",
      "rawMarkdown": "Hello @miriamredi \n\nThanks for preparing this competition and releasing it on Kaggle. The problem looks really interesting but it lacks incentive to join. A relatively large dataset that requires lots of computational power and no medals or prizes can be a red flag for lots of kagglers. This competition should at least award medals.",
      "votes": 5,
      "replies": [
        {
          "id": 1511905,
          "postDate": "2021-09-13T18:54:35.913Z",
          "content": "<p>Hi Gunes,</p>\n<p>We have to weigh multiple factors when determining whether or not to award medals, points, or prizes. In this case, as Wikipedia prides itself on being free and open to the world, it also means the competition data is possibly available. We want the competition to be a machine learning competition, instead of a data scraping competition, and removing some incentives is the only way for us to encourage such.</p>",
          "rawMarkdown": "Hi Gunes,\n\nWe have to weigh multiple factors when determining whether or not to award medals, points, or prizes. In this case, as Wikipedia prides itself on being free and open to the world, it also means the competition data is possibly available. We want the competition to be a machine learning competition, instead of a data scraping competition, and removing some incentives is the only way for us to encourage such.",
          "votes": 2
        },
        {
          "id": 1511906,
          "postDate": "2021-09-13T18:55:08.920Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1519950,
          "postDate": "2021-09-22T04:07:05.677Z",
          "content": "<p>A good incentive might be to offer selected participants - based on interesting solutions they share and open source - a possibility to co-author a research paper on the topic. </p>",
          "rawMarkdown": "A good incentive might be to offer selected participants - based on interesting solutions they share and open source - a possibility to co-author a research paper on the topic. "
        },
        {
          "id": 1610992,
          "postDate": "2021-12-07T16:41:42.153Z",
          "content": "<p>HI, I got confused by this. Are we not supposed to use data scrapping at all? I was under the impression that we were able to get, for each image url, some of the information from the wikis containing the image. Information like the title, section name, or even the description of the image (if any). Then, using ML determine which of these pieces of the information is the most relevant. </p>",
          "rawMarkdown": "HI, I got confused by this. Are we not supposed to use data scrapping at all? I was under the impression that we were able to get, for each image url, some of the information from the wikis containing the image. Information like the title, section name, or even the description of the image (if any). Then, using ML determine which of these pieces of the information is the most relevant. "
        }
      ]
    },
    {
      "id": 2087970,
      "postDate": "2023-01-06T00:34:39.657Z",
      "content": "<p>Hello, the link to get image_data_train.tar seems failed and 275GB is too large, can you give a version of .tar.gz or .zip to compress it? Or just give a successful link. Thank you!!</p>",
      "rawMarkdown": "Hello, the link to get image_data_train.tar seems failed and 275GB is too large, can you give a version of .tar.gz or .zip to compress it? Or just give a successful link. Thank you!!"
    },
    {
      "id": 1613836,
      "postDate": "2021-12-10T10:10:08.360Z",
      "content": "<p>HI <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>\n<p>Offically, it said 'The top three winning teams will receive Wikipedia-branded merchandise'.<br>\nSo where could we get the rewards? We are team \"新东方人工智能研究院\". Thanks!</p>",
      "rawMarkdown": "HI @miriamredi \n\nOffically, it said 'The top three winning teams will receive Wikipedia-branded merchandise'.\nSo where could we get the rewards? We are team \"新东方人工智能研究院\". Thanks!\n",
      "replies": [
        {
          "id": 1614031,
          "postDate": "2021-12-10T14:02:57.103Z",
          "content": "<p>Hi there! Congrats on your top position. The Kaggle team will be in touch with you in the next few days to talk about the next steps.<br>\nThanks</p>\n<p>Miriam</p>",
          "rawMarkdown": "Hi there! Congrats on your top position. The Kaggle team will be in touch with you in the next few days to talk about the next steps.\nThanks\n\nMiriam"
        }
      ]
    },
    {
      "id": 1612489,
      "postDate": "2021-12-09T02:45:41.863Z",
      "content": "<p>\"All deadlines are at 11:59 <strong>PM</strong> UTC on the corresponding day unless otherwise noted. \"<br>\nBut why the date displayed is 11:59 <strong>AM</strong> UTC?   <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>",
      "rawMarkdown": "\"All deadlines are at 11:59 **PM** UTC on the corresponding day unless otherwise noted. \"\nBut why the date displayed is 11:59 **AM** UTC?   @miriamredi ",
      "replies": [
        {
          "id": 1613168,
          "postDate": "2021-12-09T17:34:41.473Z",
          "content": "<p>Looks like a typo! Our apologies. Extended to 11:59pm UTC</p>",
          "rawMarkdown": "Looks like a typo! Our apologies. Extended to 11:59pm UTC"
        }
      ]
    },
    {
      "id": 1576243,
      "postDate": "2021-11-09T02:47:14.287Z",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>\n<p>I have a question about the test data. It seems that the resnet embeddings for test images(in image_data_test/resnet_embeddings/ directory) are only provided for about 45k images. Does this mean that only half of the image data are provided as resnet embeddings and the other half as image pixels? </p>",
      "rawMarkdown": "Hello, @miriamredi \n\nI have a question about the test data. It seems that the resnet embeddings for test images(in image_data_test/resnet_embeddings/ directory) are only provided for about 45k images. Does this mean that only half of the image data are provided as resnet embeddings and the other half as image pixels? "
    },
    {
      "id": 1513350,
      "postDate": "2021-09-15T05:14:33.517Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> and <a href=\"https://www.kaggle.com/krishnawit\" target=\"_blank\">@krishnawit</a> </p>\n<p>Thanks for hosting this challenge, it looks very interesting!</p>\n<p>I have a question about this line from competition description:</p>\n<blockquote>\n  <p>In this competition, you’ll build a model that automatically retrieves the text closest to an image.</p>\n</blockquote>\n<p>If the objective is indeed retrieval, what is the set of texts that we should retrieve from? Or is this a text generation challenge rather than retrieval? </p>",
      "rawMarkdown": "Hello @miriamredi and @krishnawit \n\nThanks for hosting this challenge, it looks very interesting!\n\nI have a question about this line from competition description:\n> In this competition, you’ll build a model that automatically retrieves the text closest to an image.\n\nIf the objective is indeed retrieval, what is the set of texts that we should retrieve from? Or is this a text generation challenge rather than retrieval? ",
      "replies": [
        {
          "id": 1513617,
          "postDate": "2021-09-15T09:20:33.170Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> -- thank you for your interested in the competition!! The objective of this competition is to predict the target <code>caption_title_and_reference_description</code> from the ones provided in the test data. You can approach this as a test generation + retrieval challenge if you prefer, but you will be evaluated based on how well you can match test images with their corresponding captions.</p>",
          "rawMarkdown": "Hi @thedrcat -- thank you for your interested in the competition!! The objective of this competition is to predict the target `caption_title_and_reference_description` from the ones provided in the test data. You can approach this as a test generation + retrieval challenge if you prefer, but you will be evaluated based on how well you can match test images with their corresponding captions.",
          "votes": 1
        },
        {
          "id": 1513851,
          "postDate": "2021-09-15T13:24:56.303Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> thanks a lot for the answer! I'm probably missing something, but I reviewed again the test data and I cannot find any file that contains <code>caption_title_and_reference_description</code>. Can you advise where specifically are these provided for test? Or should we use the ones provided in train data?</p>",
          "rawMarkdown": "Hi @miriamredi thanks a lot for the answer! I'm probably missing something, but I reviewed again the test data and I cannot find any file that contains `caption_title_and_reference_description`. Can you advise where specifically are these provided for test? Or should we use the ones provided in train data?",
          "votes": 1
        },
        {
          "id": 1522219,
          "postDate": "2021-09-24T02:00:03.110Z",
          "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> -  I think the <code>caption_title_and_reference_description</code> is in the <code>test_caption_list.csv</code> file. <img src=\"https://imgur.com/a/8P7evw5\" alt=\"\"></p>",
          "rawMarkdown": "@thedrcat -  I think the `caption_title_and_reference_description` is in the `test_caption_list.csv` file. ![](https://imgur.com/a/8P7evw5)"
        }
      ]
    },
    {
      "id": 1512039,
      "postDate": "2021-09-13T21:58:54.287Z",
      "content": "<p>Welcome to the Wikipedia Image/Caption competition!</p>\n<p>You can find more information regarding the WIT dataset in our recent SIGIR paper.<br>\n<a href=\"https://dl.acm.org/doi/abs/10.1145/3404835.3463257\" target=\"_blank\">https://dl.acm.org/doi/abs/10.1145/3404835.3463257</a></p>\n<p>Thanks for joining, we are here to help should you have any questions!</p>\n<p>Krishna.<br>\nGoogle Research</p>",
      "rawMarkdown": "Welcome to the Wikipedia Image/Caption competition!\n\nYou can find more information regarding the WIT dataset in our recent SIGIR paper.\nhttps://dl.acm.org/doi/abs/10.1145/3404835.3463257\n\nThanks for joining, we are here to help should you have any questions!\n\nKrishna.\nGoogle Research"
    }
  ],
  "comments": [
    {
      "id": 1511803,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2021-09-13T17:45:28.113000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>\n<p>Thanks for preparing this competition and releasing it on Kaggle. The problem looks really interesting but it lacks incentive to join. A relatively large dataset that requires lots of computational power and no medals or prizes can be a red flag for lots of kagglers. This competition should at least award medals.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1511905,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2021-09-13T18:54:35.913000",
          "content": "<p>Hi Gunes,</p>\n<p>We have to weigh multiple factors when determining whether or not to award medals, points, or prizes. In this case, as Wikipedia prides itself on being free and open to the world, it also means the competition data is possibly available. We want the competition to be a machine learning competition, instead of a data scraping competition, and removing some incentives is the only way for us to encourage such.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1511906,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-09-13T18:55:08.920000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1519950,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-09-22T04:07:05.677000",
          "content": "<p>A good incentive might be to offer selected participants - based on interesting solutions they share and open source - a possibility to co-author a research paper on the topic. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1610992,
          "author_name": "Mauricio CD",
          "author_url": "",
          "post_date": "2021-12-07T16:41:42.153000",
          "content": "<p>HI, I got confused by this. Are we not supposed to use data scrapping at all? I was under the impression that we were able to get, for each image url, some of the information from the wikis containing the image. Information like the title, section name, or even the description of the image (if any). Then, using ML determine which of these pieces of the information is the most relevant. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2087970,
      "author_name": "muller f",
      "author_url": "",
      "post_date": "2023-01-06T00:34:39.657000",
      "content": "<p>Hello, the link to get image_data_train.tar seems failed and 275GB is too large, can you give a version of .tar.gz or .zip to compress it? Or just give a successful link. Thank you!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1613836,
      "author_name": "TonyChen52",
      "author_url": "",
      "post_date": "2021-12-10T10:10:08.360000",
      "content": "<p>HI <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>\n<p>Offically, it said 'The top three winning teams will receive Wikipedia-branded merchandise'.<br>\nSo where could we get the rewards? We are team \"新东方人工智能研究院\". Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1614031,
          "author_name": "Miriam Redi",
          "author_url": "",
          "post_date": "2021-12-10T14:02:57.103000",
          "content": "<p>Hi there! Congrats on your top position. The Kaggle team will be in touch with you in the next few days to talk about the next steps.<br>\nThanks</p>\n<p>Miriam</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1612489,
      "author_name": "whutzy",
      "author_url": "",
      "post_date": "2021-12-09T02:45:41.863000",
      "content": "<p>\"All deadlines are at 11:59 <strong>PM</strong> UTC on the corresponding day unless otherwise noted. \"<br>\nBut why the date displayed is 11:59 <strong>AM</strong> UTC?   <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1613168,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2021-12-09T17:34:41.473000",
          "content": "<p>Looks like a typo! Our apologies. Extended to 11:59pm UTC</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1576243,
      "author_name": "Jiook Chung",
      "author_url": "",
      "post_date": "2021-11-09T02:47:14.287000",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> </p>\n<p>I have a question about the test data. It seems that the resnet embeddings for test images(in image_data_test/resnet_embeddings/ directory) are only provided for about 45k images. Does this mean that only half of the image data are provided as resnet embeddings and the other half as image pixels? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1513350,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-09-15T05:14:33.517000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> and <a href=\"https://www.kaggle.com/krishnawit\" target=\"_blank\">@krishnawit</a> </p>\n<p>Thanks for hosting this challenge, it looks very interesting!</p>\n<p>I have a question about this line from competition description:</p>\n<blockquote>\n  <p>In this competition, you’ll build a model that automatically retrieves the text closest to an image.</p>\n</blockquote>\n<p>If the objective is indeed retrieval, what is the set of texts that we should retrieve from? Or is this a text generation challenge rather than retrieval? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1513617,
          "author_name": "Miriam Redi",
          "author_url": "",
          "post_date": "2021-09-15T09:20:33.170000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> -- thank you for your interested in the competition!! The objective of this competition is to predict the target <code>caption_title_and_reference_description</code> from the ones provided in the test data. You can approach this as a test generation + retrieval challenge if you prefer, but you will be evaluated based on how well you can match test images with their corresponding captions.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1513851,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-09-15T13:24:56.303000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/miriamredi\" target=\"_blank\">@miriamredi</a> thanks a lot for the answer! I'm probably missing something, but I reviewed again the test data and I cannot find any file that contains <code>caption_title_and_reference_description</code>. Can you advise where specifically are these provided for test? Or should we use the ones provided in train data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1522219,
          "author_name": "Callistus Ndemo",
          "author_url": "",
          "post_date": "2021-09-24T02:00:03.110000",
          "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> -  I think the <code>caption_title_and_reference_description</code> is in the <code>test_caption_list.csv</code> file. <img src=\"https://imgur.com/a/8P7evw5\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1512039,
      "author_name": "Krishna Srinivasan",
      "author_url": "",
      "post_date": "2021-09-13T21:58:54.287000",
      "content": "<p>Welcome to the Wikipedia Image/Caption competition!</p>\n<p>You can find more information regarding the WIT dataset in our recent SIGIR paper.<br>\n<a href=\"https://dl.acm.org/doi/abs/10.1145/3404835.3463257\" target=\"_blank\">https://dl.acm.org/doi/abs/10.1145/3404835.3463257</a></p>\n<p>Thanks for joining, we are here to help should you have any questions!</p>\n<p>Krishna.<br>\nGoogle Research</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1511796": "Welcome to the Wikipedia Image/Caption competition!\n\nWe are beyond excited to support this competition, and call for brilliant people like you to solve challenging problems in the vision and language space! In this challenge, you will  play with large amounts of multilingual and multimodal data, coming from the world's largest online encyclopedia. \n\nYour contributions to this competition are contributions to free and open knowledge!\n\nThanks for joining, we are here to help should you have any questions!\n\nMiriam\n",
    "1511803": "Hello @miriamredi \n\nThanks for preparing this competition and releasing it on Kaggle. The problem looks really interesting but it lacks incentive to join. A relatively large dataset that requires lots of computational power and no medals or prizes can be a red flag for lots of kagglers. This competition should at least award medals.",
    "2087970": "Hello, the link to get image_data_train.tar seems failed and 275GB is too large, can you give a version of .tar.gz or .zip to compress it? Or just give a successful link. Thank you!!",
    "1613836": "HI @miriamredi \n\nOffically, it said 'The top three winning teams will receive Wikipedia-branded merchandise'.\nSo where could we get the rewards? We are team \"新东方人工智能研究院\". Thanks!\n",
    "1612489": "\"All deadlines are at 11:59 **PM** UTC on the corresponding day unless otherwise noted. \"\nBut why the date displayed is 11:59 **AM** UTC?   @miriamredi ",
    "1576243": "Hello, @miriamredi \n\nI have a question about the test data. It seems that the resnet embeddings for test images(in image_data_test/resnet_embeddings/ directory) are only provided for about 45k images. Does this mean that only half of the image data are provided as resnet embeddings and the other half as image pixels? ",
    "1513350": "Hello @miriamredi and @krishnawit \n\nThanks for hosting this challenge, it looks very interesting!\n\nI have a question about this line from competition description:\n> In this competition, you’ll build a model that automatically retrieves the text closest to an image.\n\nIf the objective is indeed retrieval, what is the set of texts that we should retrieve from? Or is this a text generation challenge rather than retrieval? ",
    "1512039": "Welcome to the Wikipedia Image/Caption competition!\n\nYou can find more information regarding the WIT dataset in our recent SIGIR paper.\nhttps://dl.acm.org/doi/abs/10.1145/3404835.3463257\n\nThanks for joining, we are here to help should you have any questions!\n\nKrishna.\nGoogle Research"
  }
}