{
  "id": 94410,
  "title": "A guess of \"leak\" detection via Google Search",
  "url": "/competitions/imet-2019-fgvc6/discussion/94410",
  "author_name": "Yiheng Wang",
  "post_date": "2019-06-04T10:25:51.933000",
  "votes": 17,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I use google image search for a few images from the training set, and the return words may contain a part of labels, or their synonyms. My guess is that people can use this kind of function to retrieve all test images (although it is only useful in stage 1, they can make use of these pseudo labels for knowledge distillation or augment their training set), and then get each image's return word. Utilizing word vector can help them match the closest labels.\nIn the following image one, the word chalice has the label 533, and the word french is 147. Both of them are true. </p>\n\n<p>Certainly, it is just my simple trial and get the label directly is still very hard in this way, at the same time, I promise that I only searched ~15 images from the training set, nothing will be changed for our solutions.  (!!! <strong>even if it is a 1 stage competition, searching for test set is not allowed!!!</strong>)</p>\n\n<p>Wish we all can achieve a reasonable result in stage 2, and wish to know what top participants really did for their decent results.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13392/e1.png\" alt=\"e1\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13393/e2.png\" alt=\"e2\"></p>",
  "messages": [
    {
      "id": 543094,
      "postDate": "2019-06-04T10:25:51.933Z",
      "content": "<p>I use google image search for a few images from the training set, and the return words may contain a part of labels, or their synonyms. My guess is that people can use this kind of function to retrieve all test images (although it is only useful in stage 1, they can make use of these pseudo labels for knowledge distillation or augment their training set), and then get each image's return word. Utilizing word vector can help them match the closest labels.\nIn the following image one, the word chalice has the label 533, and the word french is 147. Both of them are true. </p>\n\n<p>Certainly, it is just my simple trial and get the label directly is still very hard in this way, at the same time, I promise that I only searched ~15 images from the training set, nothing will be changed for our solutions.  (!!! <strong>even if it is a 1 stage competition, searching for test set is not allowed!!!</strong>)</p>\n\n<p>Wish we all can achieve a reasonable result in stage 2, and wish to know what top participants really did for their decent results.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13392/e1.png\" alt=\"e1\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13393/e2.png\" alt=\"e2\"></p>",
      "rawMarkdown": "I use google image search for a few images from the training set, and the return words may contain a part of labels, or their synonyms. My guess is that people can use this kind of function to retrieve all test images (although it is only useful in stage 1, they can make use of these pseudo labels for knowledge distillation or augment their training set), and then get each image's return word. Utilizing word vector can help them match the closest labels.\nIn the following image one, the word chalice has the label 533, and the word french is 147. Both of them are true. \n\nCertainly, it is just my simple trial and get the label directly is still very hard in this way, at the same time, I promise that I only searched ~15 images from the training set, nothing will be changed for our solutions.  (!!! **even if it is a 1 stage competition, searching for test set is not allowed!!!**)\n\nWish we all can achieve a reasonable result in stage 2, and wish to know what top participants really did for their decent results.\n\n![e1](https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13392/e1.png)\n\n![e2](https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13393/e2.png)",
      "votes": 17
    },
    {
      "id": 543253,
      "postDate": "2019-06-04T12:19:33.137Z",
      "content": "<p>I guess this helps a lot. I used a trained model to predict missing labels on train images and use that soft labels to train a new model. This improves LB. </p>",
      "rawMarkdown": "I guess this helps a lot. I used a trained model to predict missing labels on train images and use that soft labels to train a new model. This improves LB. ",
      "votes": 7,
      "replies": [
        {
          "id": 543283,
          "postDate": "2019-06-04T12:47:18.337Z",
          "content": "<p>How much did it improved LB?</p>",
          "rawMarkdown": "How much did it improved LB?",
          "votes": 1
        },
        {
          "id": 543284,
          "postDate": "2019-06-04T12:50:36.120Z",
          "content": "<p>LB 0.652 to 0.662 (6-folds se_resnext101)</p>",
          "rawMarkdown": "LB 0.652 to 0.662 (6-folds se_resnext101)",
          "votes": 9
        },
        {
          "id": 544197,
          "postDate": "2019-06-05T08:57:51.657Z",
          "content": "<p>nice improvement, the only pity is that we do not have time to implement this idea and train another model.</p>",
          "rawMarkdown": "nice improvement, the only pity is that we do not have time to implement this idea and train another model."
        }
      ]
    },
    {
      "id": 543203,
      "postDate": "2019-06-04T11:26:39.257Z",
      "content": "<p>Actually you can find the 'correct' label directly. </p>\n\n<p>As mentioned by <a href=\"https://www.kaggle.com/tereka\">tereka</a>, the data are taken from <a href=\"https://www.metmuseum.org/art/collection.\">The Metropolitan Muesum of Art in New York</a>. I used google search with several images in the train and test set and found them on the website. </p>\n\n<p>For example, '1008c7837081f985.png' in the train-set, you can find it <a href=\"https://www.metmuseum.org/art/collection/search/42324\">here</a>. In the Object Details, you can find: 'Culture: China', 'Object Type/Material: Ceramics, Pottery, Reliefs, Sculpture, Stoneware, Vases, Vessels'. If you check the labels of this image in the train set, they are: 79 -&gt; 'culture::china' and 1062 -&gt; 'tag::vases'. </p>\n\n<p>Another example for the test, '1a529c44ca7c4209.png', you can find it <a href=\"https://www.metmuseum.org/art/collection/search/36674\">here</a>. Culture and Object Type/Material information can also be found on the page. </p>\n\n<p>With this information in hand, you don't need to clean train.csv. If one found the relation between the image id in train/test set (1008c7837081f985) and on the website (42324), he can use a crawler to get all the information from Metropolitan Muesum website and use it as a reference. The next step is simply replace the wrong label predicted by the model with the label crawled from the website. If you save the crawled labels in a .csv, there is no need to have internet connection. If the crawled label .csv is provided, what is the time you need to create a significant boosted solution based on one that you already have? I think experienced kagglers already have the answer. </p>\n\n<p>I haven't done it and don't plan to do it. I believe it violates the rule that 'Besides using the provided dataset, participants are restricted from collecting additional data for this competition. '  </p>",
      "rawMarkdown": "Actually you can find the 'correct' label directly. \n\nAs mentioned by [tereka](https://www.kaggle.com/tereka), the data are taken from [The Metropolitan Muesum of Art in New York](https://www.metmuseum.org/art/collection.). I used google search with several images in the train and test set and found them on the website. \n\nFor example, '1008c7837081f985.png' in the train-set, you can find it [here](https://www.metmuseum.org/art/collection/search/42324). In the Object Details, you can find: 'Culture: China', 'Object Type/Material: Ceramics, Pottery, Reliefs, Sculpture, Stoneware, Vases, Vessels'. If you check the labels of this image in the train set, they are: 79 -&gt; 'culture::china' and 1062 -&gt; 'tag::vases'. \n\nAnother example for the test, '1a529c44ca7c4209.png', you can find it [here](https://www.metmuseum.org/art/collection/search/36674). Culture and Object Type/Material information can also be found on the page. \n\nWith this information in hand, you don't need to clean train.csv. If one found the relation between the image id in train/test set (1008c7837081f985) and on the website (42324), he can use a crawler to get all the information from Metropolitan Muesum website and use it as a reference. The next step is simply replace the wrong label predicted by the model with the label crawled from the website. If you save the crawled labels in a .csv, there is no need to have internet connection. If the crawled label .csv is provided, what is the time you need to create a significant boosted solution based on one that you already have? I think experienced kagglers already have the answer. \n\nI haven't done it and don't plan to do it. I believe it violates the rule that 'Besides using the provided dataset, participants are restricted from collecting additional data for this competition. '  ",
      "votes": 6
    },
    {
      "id": 543138,
      "postDate": "2019-06-04T10:49:25.310Z",
      "content": "<p>This won't help on stage 2</p>",
      "rawMarkdown": "This won't help on stage 2",
      "votes": 1,
      "replies": [
        {
          "id": 543147,
          "postDate": "2019-06-04T10:56:00.743Z",
          "content": "<p>But you can use the extra labeled data to train your model. This affects the fairness of the competition.</p>",
          "rawMarkdown": "But you can use the extra labeled data to train your model. This affects the fairness of the competition.",
          "votes": 3
        },
        {
          "id": 543162,
          "postDate": "2019-06-04T11:04:10.970Z",
          "content": "<p>But this is against the rules?</p>\n\n<p>-No internet access enabled\n-Only whitelisted data is allowed</p>",
          "rawMarkdown": "But this is against the rules?\n\n-No internet access enabled\n-Only whitelisted data is allowed"
        },
        {
          "id": 543217,
          "postDate": "2019-06-04T11:41:40.860Z",
          "content": "<p>Yes, We cannot guarantee that everybody follows the rules.</p>",
          "rawMarkdown": "Yes, We cannot guarantee that everybody follows the rules."
        }
      ]
    },
    {
      "id": 543175,
      "postDate": "2019-06-04T11:14:11.953Z",
      "content": "<p>It would make a lot of sense to clean train.csv this way. It has <strong>a lot</strong> of missing labels and it also has some labels that are clearly wrong.</p>",
      "rawMarkdown": "It would make a lot of sense to clean train.csv this way. It has **a lot** of missing labels and it also has some labels that are clearly wrong.",
      "votes": 2,
      "replies": [
        {
          "id": 543224,
          "postDate": "2019-06-04T11:50:27.217Z",
          "content": "<p>If you check the competition github page <a href=\"https://github.com/visipedia/imet-fgvcx\">(link)</a>, the labels are meant to be noisy. </p>",
          "rawMarkdown": "If you check the competition github page [(link)](https://github.com/visipedia/imet-fgvcx), the labels are meant to be noisy. "
        }
      ]
    },
    {
      "id": 543286,
      "postDate": "2019-06-04T12:51:12.963Z",
      "content": "<p>Moreover, you share only kernel code with weights, no way of obtaining information about using with crawled data.</p>",
      "rawMarkdown": "Moreover, you share only kernel code with weights, no way of obtaining information about using with crawled data.",
      "replies": [
        {
          "id": 543381,
          "postDate": "2019-06-04T14:02:28.490Z",
          "content": "<p>thus for top participants, like top 11, I suggest that they should reproduce their training result to prove themselves.</p>",
          "rawMarkdown": "thus for top participants, like top 11, I suggest that they should reproduce their training result to prove themselves."
        },
        {
          "id": 544195,
          "postDate": "2019-06-05T08:57:02.850Z",
          "content": "<p>I agree, even though there are some randomness, the results should not vary much..</p>",
          "rawMarkdown": "I agree, even though there are some randomness, the results should not vary much.."
        }
      ]
    },
    {
      "id": 543112,
      "postDate": "2019-06-04T10:36:02.050Z",
      "content": "<p>Terrible discovery! We cannot guarantee that everybody follows the rules and don't use these extra images/labels.</p>",
      "rawMarkdown": "Terrible discovery! We cannot guarantee that everybody follows the rules and don't use these extra images/labels.",
      "replies": [
        {
          "id": 543130,
          "postDate": "2019-06-04T10:45:27.160Z",
          "content": "<p>yes, thus we need 2 stages, and submit the codes. </p>",
          "rawMarkdown": "yes, thus we need 2 stages, and submit the codes. "
        }
      ]
    },
    {
      "id": 543132,
      "postDate": "2019-06-04T10:47:09.670Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 543253,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2019-06-04T12:19:33.137000",
      "content": "<p>I guess this helps a lot. I used a trained model to predict missing labels on train images and use that soft labels to train a new model. This improves LB. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 543283,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2019-06-04T12:47:18.337000",
          "content": "<p>How much did it improved LB?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543284,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2019-06-04T12:50:36.120000",
          "content": "<p>LB 0.652 to 0.662 (6-folds se_resnext101)</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 544197,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2019-06-05T08:57:51.657000",
          "content": "<p>nice improvement, the only pity is that we do not have time to implement this idea and train another model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543203,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2019-06-04T11:26:39.257000",
      "content": "<p>Actually you can find the 'correct' label directly. </p>\n\n<p>As mentioned by <a href=\"https://www.kaggle.com/tereka\">tereka</a>, the data are taken from <a href=\"https://www.metmuseum.org/art/collection.\">The Metropolitan Muesum of Art in New York</a>. I used google search with several images in the train and test set and found them on the website. </p>\n\n<p>For example, '1008c7837081f985.png' in the train-set, you can find it <a href=\"https://www.metmuseum.org/art/collection/search/42324\">here</a>. In the Object Details, you can find: 'Culture: China', 'Object Type/Material: Ceramics, Pottery, Reliefs, Sculpture, Stoneware, Vases, Vessels'. If you check the labels of this image in the train set, they are: 79 -&gt; 'culture::china' and 1062 -&gt; 'tag::vases'. </p>\n\n<p>Another example for the test, '1a529c44ca7c4209.png', you can find it <a href=\"https://www.metmuseum.org/art/collection/search/36674\">here</a>. Culture and Object Type/Material information can also be found on the page. </p>\n\n<p>With this information in hand, you don't need to clean train.csv. If one found the relation between the image id in train/test set (1008c7837081f985) and on the website (42324), he can use a crawler to get all the information from Metropolitan Muesum website and use it as a reference. The next step is simply replace the wrong label predicted by the model with the label crawled from the website. If you save the crawled labels in a .csv, there is no need to have internet connection. If the crawled label .csv is provided, what is the time you need to create a significant boosted solution based on one that you already have? I think experienced kagglers already have the answer. </p>\n\n<p>I haven't done it and don't plan to do it. I believe it violates the rule that 'Besides using the provided dataset, participants are restricted from collecting additional data for this competition. '  </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 543138,
      "author_name": "Anna Novikova",
      "author_url": "",
      "post_date": "2019-06-04T10:49:25.310000",
      "content": "<p>This won't help on stage 2</p>",
      "votes": 1,
      "replies": [
        {
          "id": 543147,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2019-06-04T10:56:00.743000",
          "content": "<p>But you can use the extra labeled data to train your model. This affects the fairness of the competition.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 543162,
          "author_name": "Anna Novikova",
          "author_url": "",
          "post_date": "2019-06-04T11:04:10.970000",
          "content": "<p>But this is against the rules?</p>\n\n<p>-No internet access enabled\n-Only whitelisted data is allowed</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543217,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2019-06-04T11:41:40.860000",
          "content": "<p>Yes, We cannot guarantee that everybody follows the rules.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543175,
      "author_name": "Artyom Palvelev",
      "author_url": "",
      "post_date": "2019-06-04T11:14:11.953000",
      "content": "<p>It would make a lot of sense to clean train.csv this way. It has <strong>a lot</strong> of missing labels and it also has some labels that are clearly wrong.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 543224,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-06-04T11:50:27.217000",
          "content": "<p>If you check the competition github page <a href=\"https://github.com/visipedia/imet-fgvcx\">(link)</a>, the labels are meant to be noisy. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543286,
      "author_name": "DmitryKustikov",
      "author_url": "",
      "post_date": "2019-06-04T12:51:12.963000",
      "content": "<p>Moreover, you share only kernel code with weights, no way of obtaining information about using with crawled data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 543381,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2019-06-04T14:02:28.490000",
          "content": "<p>thus for top participants, like top 11, I suggest that they should reproduce their training result to prove themselves.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 544195,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2019-06-05T08:57:02.850000",
          "content": "<p>I agree, even though there are some randomness, the results should not vary much..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543112,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-06-04T10:36:02.050000",
      "content": "<p>Terrible discovery! We cannot guarantee that everybody follows the rules and don't use these extra images/labels.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 543130,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2019-06-04T10:45:27.160000",
          "content": "<p>yes, thus we need 2 stages, and submit the codes. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543132,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T10:47:09.670000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543094": "I use google image search for a few images from the training set, and the return words may contain a part of labels, or their synonyms. My guess is that people can use this kind of function to retrieve all test images (although it is only useful in stage 1, they can make use of these pseudo labels for knowledge distillation or augment their training set), and then get each image's return word. Utilizing word vector can help them match the closest labels.\nIn the following image one, the word chalice has the label 533, and the word french is 147. Both of them are true. \n\nCertainly, it is just my simple trial and get the label directly is still very hard in this way, at the same time, I promise that I only searched ~15 images from the training set, nothing will be changed for our solutions.  (!!! **even if it is a 1 stage competition, searching for test set is not allowed!!!**)\n\nWish we all can achieve a reasonable result in stage 2, and wish to know what top participants really did for their decent results.\n\n![e1](https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13392/e1.png)\n\n![e2](https://storage.googleapis.com/kaggle-forum-message-attachments/543094/13393/e2.png)",
    "543253": "I guess this helps a lot. I used a trained model to predict missing labels on train images and use that soft labels to train a new model. This improves LB. ",
    "543203": "Actually you can find the 'correct' label directly. \n\nAs mentioned by [tereka](https://www.kaggle.com/tereka), the data are taken from [The Metropolitan Muesum of Art in New York](https://www.metmuseum.org/art/collection.). I used google search with several images in the train and test set and found them on the website. \n\nFor example, '1008c7837081f985.png' in the train-set, you can find it [here](https://www.metmuseum.org/art/collection/search/42324). In the Object Details, you can find: 'Culture: China', 'Object Type/Material: Ceramics, Pottery, Reliefs, Sculpture, Stoneware, Vases, Vessels'. If you check the labels of this image in the train set, they are: 79 -&gt; 'culture::china' and 1062 -&gt; 'tag::vases'. \n\nAnother example for the test, '1a529c44ca7c4209.png', you can find it [here](https://www.metmuseum.org/art/collection/search/36674). Culture and Object Type/Material information can also be found on the page. \n\nWith this information in hand, you don't need to clean train.csv. If one found the relation between the image id in train/test set (1008c7837081f985) and on the website (42324), he can use a crawler to get all the information from Metropolitan Muesum website and use it as a reference. The next step is simply replace the wrong label predicted by the model with the label crawled from the website. If you save the crawled labels in a .csv, there is no need to have internet connection. If the crawled label .csv is provided, what is the time you need to create a significant boosted solution based on one that you already have? I think experienced kagglers already have the answer. \n\nI haven't done it and don't plan to do it. I believe it violates the rule that 'Besides using the provided dataset, participants are restricted from collecting additional data for this competition. '  ",
    "543138": "This won't help on stage 2",
    "543175": "It would make a lot of sense to clean train.csv this way. It has **a lot** of missing labels and it also has some labels that are clearly wrong.",
    "543286": "Moreover, you share only kernel code with weights, no way of obtaining information about using with crawled data.",
    "543112": "Terrible discovery! We cannot guarantee that everybody follows the rules and don't use these extra images/labels.",
    "543132": ""
  }
}