{
  "id": 286451,
  "title": "Major inconvenience using binary classification for this project: Making submission itself seems too hard",
  "url": "/competitions/wikipedia-image-caption/discussion/286451",
  "author_name": "",
  "post_date": "2021-11-09T07:36:29.090628500Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am going for an approach that classifies the image-caption pair as wrong(0) or right(1). I used negative sampling to generate the wrong pairs and fed them all to the model. The accuracy was about 0.7 which is still very bad, but I thought that it was pretty good for a first try, regarding the simplicity of the model.</p>\n<p>However, at the time I tried to make a submission, I just realized that I needed to match 90k captions with 90k images. This means that predicting the most likely caption for a single image takes 90k image-caption pairs to be fed to the model, and predicting it for all the images takes 90k*90k=<strong>8.1b image-caption pairs</strong>. It seems almost impossible to go through all 8.1b of them. I tried to save time for the operation but it took at least 10 seconds to predict for a single image. This means that the whole process will take about 250 hours to go through.</p>\n<p>Therefore I am now thinking that binary classification (at least using it for a single image-single caption pair) just may not suit this problem well. Maybe there could be another approach that makes matching lots of pairs much easier. Maybe my code is just too inefficient! </p>\n<p>I am quite new for machine learning, and this is by far the most challenging project for me. It would be a pleasure to hear your thoughts and experience about this problem.</p>",
  "messages": [
    {
      "id": "1576450",
      "postDate": "11/09/2021 07:36:29",
      "content": "<p>I am going for an approach that classifies the image-caption pair as wrong(0) or right(1). I used negative sampling to generate the wrong pairs and fed them all to the model. The accuracy was about 0.7 which is still very bad, but I thought that it was pretty good for a first try, regarding the simplicity of the model.</p>\n<p>However, at the time I tried to make a submission, I just realized that I needed to match 90k captions with 90k images. This means that predicting the most likely caption for a single image takes 90k image-caption pairs to be fed to the model, and predicting it for all the images takes 90k*90k=<strong>8.1b image-caption pairs</strong>. It seems almost impossible to go through all 8.1b of them. I tried to save time for the operation but it took at least 10 seconds to predict for a single image. This means that the whole process will take about 250 hours to go through.</p>\n<p>Therefore I am now thinking that binary classification (at least using it for a single image-single caption pair) just may not suit this problem well. Maybe there could be another approach that makes matching lots of pairs much easier. Maybe my code is just too inefficient! </p>\n<p>I am quite new for machine learning, and this is by far the most challenging project for me. It would be a pleasure to hear your thoughts and experience about this problem.</p>",
      "rawMarkdown": "I am going for an approach that classifies the image-caption pair as wrong(0) or right(1). I used negative sampling to generate the wrong pairs and fed them all to the model. The accuracy was about 0.7 which is still very bad, but I thought that it was pretty good for a first try, regarding the simplicity of the model.\n\nHowever, at the time I tried to make a submission, I just realized that I needed to match 90k captions with 90k images. This means that predicting the most likely caption for a single image takes 90k image-caption pairs to be fed to the model, and predicting it for all the images takes 90k*90k=**8.1b image-caption pairs**. It seems almost impossible to go through all 8.1b of them. I tried to save time for the operation but it took at least 10 seconds to predict for a single image. This means that the whole process will take about 250 hours to go through.\n\nTherefore I am now thinking that binary classification (at least using it for a single image-single caption pair) just may not suit this problem well. Maybe there could be another approach that makes matching lots of pairs much easier. Maybe my code is just too inefficient! \n\nI am quite new for machine learning, and this is by far the most challenging project for me. It would be a pleasure to hear your thoughts and experience about this problem.",
      "votes": null
    },
    {
      "id": "1576652",
      "postDate": "11/09/2021 11:18:15",
      "content": "<p>You can compute the text embeddings before hand and save<br>\nThen you can use <a href=\"https://github.com/facebookresearch/faiss\" target=\"_blank\">Faiss</a> to do similarity search<br>\nI haven't tried it personally but have heard that it is fast</p>",
      "rawMarkdown": "You can compute the text embeddings before hand and save\nThen you can use [Faiss](https://github.com/facebookresearch/faiss) to do similarity search\nI haven't tried it personally but have heard that it is fast",
      "votes": null
    },
    {
      "id": "1577310",
      "postDate": "11/10/2021 01:15:19",
      "content": "<p>Hello again! <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> <br>\nI have thought of precomputing the embedding before hand, but I thought I should still calculate for all the pairs in a brute force style. Faiss seems to be really helpful for this kind of searching! Thank you for sharing your knowledge. I really appreciate it!</p>",
      "rawMarkdown": "Hello again! @debarshichanda \nI have thought of precomputing the embedding before hand, but I thought I should still calculate for all the pairs in a brute force style. Faiss seems to be really helpful for this kind of searching! Thank you for sharing your knowledge. I really appreciate it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1576652,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "11/09/2021 11:18:15",
      "content": "<p>You can compute the text embeddings before hand and save<br>\nThen you can use <a href=\"https://github.com/facebookresearch/faiss\" target=\"_blank\">Faiss</a> to do similarity search<br>\nI haven't tried it personally but have heard that it is fast</p>",
      "votes": null,
      "replies": [
        {
          "id": 1577310,
          "author_name": "jiookchung",
          "author_url": "",
          "post_date": "11/10/2021 01:15:19",
          "content": "<p>Hello again! <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> <br>\nI have thought of precomputing the embedding before hand, but I thought I should still calculate for all the pairs in a brute force style. Faiss seems to be really helpful for this kind of searching! Thank you for sharing your knowledge. I really appreciate it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1576450": "I am going for an approach that classifies the image-caption pair as wrong(0) or right(1). I used negative sampling to generate the wrong pairs and fed them all to the model. The accuracy was about 0.7 which is still very bad, but I thought that it was pretty good for a first try, regarding the simplicity of the model.\n\nHowever, at the time I tried to make a submission, I just realized that I needed to match 90k captions with 90k images. This means that predicting the most likely caption for a single image takes 90k image-caption pairs to be fed to the model, and predicting it for all the images takes 90k*90k=**8.1b image-caption pairs**. It seems almost impossible to go through all 8.1b of them. I tried to save time for the operation but it took at least 10 seconds to predict for a single image. This means that the whole process will take about 250 hours to go through.\n\nTherefore I am now thinking that binary classification (at least using it for a single image-single caption pair) just may not suit this problem well. Maybe there could be another approach that makes matching lots of pairs much easier. Maybe my code is just too inefficient! \n\nI am quite new for machine learning, and this is by far the most challenging project for me. It would be a pleasure to hear your thoughts and experience about this problem.",
    "1576652": "You can compute the text embeddings before hand and save\nThen you can use [Faiss](https://github.com/facebookresearch/faiss) to do similarity search\nI haven't tried it personally but have heard that it is fast",
    "1577310": "Hello again! @debarshichanda \nI have thought of precomputing the embedding before hand, but I thought I should still calculate for all the pairs in a brute force style. Faiss seems to be really helpful for this kind of searching! Thank you for sharing your knowledge. I really appreciate it!"
  },
  "source": "meta"
}