{
  "id": 272846,
  "title": "Possible error with test embeddings",
  "url": "/competitions/wikipedia-image-caption/discussion/272846",
  "author_name": "",
  "post_date": "2021-09-17T17:26:11.291499Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Unless I am mistaken, the test data has embeddings of size 2047 and the train data has a size of 2048. The competition lists the test embeddings to be of size 2048. </p>\n<p>Is this the correct training data?<br>\n<a href=\"https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/\" target=\"_blank\">https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/</a></p>",
  "messages": [
    {
      "id": "1515937",
      "postDate": "09/17/2021 17:26:11",
      "content": "<p>Unless I am mistaken, the test data has embeddings of size 2047 and the train data has a size of 2048. The competition lists the test embeddings to be of size 2048. </p>\n<p>Is this the correct training data?<br>\n<a href=\"https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/\" target=\"_blank\">https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/</a></p>",
      "rawMarkdown": "Unless I am mistaken, the test data has embeddings of size 2047 and the train data has a size of 2048. The competition lists the test embeddings to be of size 2048. \n\nIs this the correct training data?\nhttps://analytics.wikimedia.org/published/datasets/one-off/caption_competition/",
      "votes": null
    },
    {
      "id": "1518346",
      "postDate": "09/20/2021 15:21:38",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcmiller\" target=\"_blank\">@pcmiller</a> for reaching out. Yes - what you shared is the correct data link! Kaggle has also stored a copy of the data here: <a href=\"https://storage.cloud.google.com/wikimedia-image-caption-public/image_data_train.tar\" target=\"_blank\">https://storage.cloud.google.com/wikimedia-image-caption-public/image_data_train.tar</a></p>\n<p>We checked  both the training and the test_val embeddings, and all the embeddings in these files are 2048 long. Could you please post some example snippets of data or code so that we can check where the error is?</p>\n<p>Thanks!</p>\n<p>Miriam</p>",
      "rawMarkdown": "Thanks @pcmiller for reaching out. Yes - what you shared is the correct data link! Kaggle has also stored a copy of the data here: https://storage.cloud.google.com/wikimedia-image-caption-public/image_data_train.tar\n\n We checked  both the training and the test_val embeddings, and all the embeddings in these files are 2048 long. Could you please post some example snippets of data or code so that we can check where the error is?\n\nThanks!\n\nMiriam",
      "votes": null
    },
    {
      "id": "1523636",
      "postDate": "09/25/2021 15:22:01",
      "content": "<p>I was also confused by this since the previews from Kaggle show that there are 2048 columns (but 1 is for the url).</p>\n<p>However, at a closer look of the urls we can see that the first embedding column is smooshed in together (because of a different delimiter), so we need to read them this way:</p>\n<p><code>pd.read_csv(fn, header=None, sep=\"\\s+|,\")</code></p>\n<p>then we can see the files from both sources are indeed with 2048 columns of embeddings. </p>",
      "rawMarkdown": "I was also confused by this since the previews from Kaggle show that there are 2048 columns (but 1 is for the url).\n\nHowever, at a closer look of the urls we can see that the first embedding column is smooshed in together (because of a different delimiter), so we need to read them this way:\n\n`pd.read_csv(fn, header=None, sep=\"\\s+|,\")`\n\nthen we can see the files from both sources are indeed with 2048 columns of embeddings.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1518346,
      "author_name": "miriamredi",
      "author_url": "",
      "post_date": "09/20/2021 15:21:38",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcmiller\" target=\"_blank\">@pcmiller</a> for reaching out. Yes - what you shared is the correct data link! Kaggle has also stored a copy of the data here: <a href=\"https://storage.cloud.google.com/wikimedia-image-caption-public/image_data_train.tar\" target=\"_blank\">https://storage.cloud.google.com/wikimedia-image-caption-public/image_data_train.tar</a></p>\n<p>We checked  both the training and the test_val embeddings, and all the embeddings in these files are 2048 long. Could you please post some example snippets of data or code so that we can check where the error is?</p>\n<p>Thanks!</p>\n<p>Miriam</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1523636,
      "author_name": "b0ww1d",
      "author_url": "",
      "post_date": "09/25/2021 15:22:01",
      "content": "<p>I was also confused by this since the previews from Kaggle show that there are 2048 columns (but 1 is for the url).</p>\n<p>However, at a closer look of the urls we can see that the first embedding column is smooshed in together (because of a different delimiter), so we need to read them this way:</p>\n<p><code>pd.read_csv(fn, header=None, sep=\"\\s+|,\")</code></p>\n<p>then we can see the files from both sources are indeed with 2048 columns of embeddings. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1515937": "Unless I am mistaken, the test data has embeddings of size 2047 and the train data has a size of 2048. The competition lists the test embeddings to be of size 2048. \n\nIs this the correct training data?\nhttps://analytics.wikimedia.org/published/datasets/one-off/caption_competition/",
    "1518346": "Thanks @pcmiller for reaching out. Yes - what you shared is the correct data link! Kaggle has also stored a copy of the data here: https://storage.cloud.google.com/wikimedia-image-caption-public/image_data_train.tar\n\n We checked  both the training and the test_val embeddings, and all the embeddings in these files are 2048 long. Could you please post some example snippets of data or code so that we can check where the error is?\n\nThanks!\n\nMiriam",
    "1523636": "I was also confused by this since the previews from Kaggle show that there are 2048 columns (but 1 is for the url).\n\nHowever, at a closer look of the urls we can see that the first embedding column is smooshed in together (because of a different delimiter), so we need to read them this way:\n\n`pd.read_csv(fn, header=None, sep=\"\\s+|,\")`\n\nthen we can see the files from both sources are indeed with 2048 columns of embeddings."
  },
  "source": "meta"
}