{
  "id": 284720,
  "title": "Easy-to-Use Dataset",
  "url": "/competitions/wikipedia-image-caption/discussion/284720",
  "author_name": "",
  "post_date": "2021-11-02T06:03:15.688196300Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Dataset Creation and Versioning: <a href=\"https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset\" target=\"_blank\">https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset</a><br>\n</p>\n<p>Some key points to notice are that not all the images have captions corresponding to them, and those samples have been filtered from this dataset.<br>\n</p><hr><br>\nWeights and Biases has been used for Dataset Versioning<br>\nLink: <a href=\"https://wandb.ai/dchanda/Wikipedia/artifacts/dataset/Wiki-data/ac20734be4747897b5ba\" target=\"_blank\">WandB Artifact</a><br>\n<img src=\"https://i.imgur.com/3neB8pj.jpg\" alt=\"\"><br>\n<hr><br>\nThe latest dataset can be downloaded using the following code:<p></p>\n<pre><code>run = wandb.init(project=\"Wikipedia\", \n                 anonymous=\"must\")\nartifact = run.use_artifact('dchanda/Wikipedia/Wiki-data:latest', type='dataset')\nartifact_dir = artifact.download()\nrun.finish()\n\nfor file in os.listdir(artifact_dir):\n    filepath = os.path.join(artifact_dir, file)\n    with open(filepath, \"rb\") as fp:\n        contents = pickle.load(fp)\n</code></pre>\n<p></p><hr><br>\nV0: Contains samples from file 00000-00003 from the dataset provided at  <a href=\"https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/\" target=\"_blank\">analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/</a><p></p>\n<p>Will add samples from more files in the corresponding versions</p>",
  "messages": [
    {
      "id": "1567682",
      "postDate": "11/02/2021 06:03:15",
      "content": "<p>Dataset Creation and Versioning: <a href=\"https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset\" target=\"_blank\">https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset</a><br>\n</p>\n<p>Some key points to notice are that not all the images have captions corresponding to them, and those samples have been filtered from this dataset.<br>\n</p><hr><br>\nWeights and Biases has been used for Dataset Versioning<br>\nLink: <a href=\"https://wandb.ai/dchanda/Wikipedia/artifacts/dataset/Wiki-data/ac20734be4747897b5ba\" target=\"_blank\">WandB Artifact</a><br>\n<img src=\"https://i.imgur.com/3neB8pj.jpg\" alt=\"\"><br>\n<hr><br>\nThe latest dataset can be downloaded using the following code:<p></p>\n<pre><code>run = wandb.init(project=\"Wikipedia\", \n                 anonymous=\"must\")\nartifact = run.use_artifact('dchanda/Wikipedia/Wiki-data:latest', type='dataset')\nartifact_dir = artifact.download()\nrun.finish()\n\nfor file in os.listdir(artifact_dir):\n    filepath = os.path.join(artifact_dir, file)\n    with open(filepath, \"rb\") as fp:\n        contents = pickle.load(fp)\n</code></pre>\n<p></p><hr><br>\nV0: Contains samples from file 00000-00003 from the dataset provided at  <a href=\"https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/\" target=\"_blank\">analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/</a><p></p>\n<p>Will add samples from more files in the corresponding versions</p>",
      "rawMarkdown": "Dataset Creation and Versioning: [https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset](https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset)\n<span>This dataset contains 300px <code>b64_bytes</code> for all the images (no downloading required) along with a <code>list of captions</code> in all the available languages</span>\n\nSome key points to notice are that not all the images have captions corresponding to them, and those samples have been filtered from this dataset.\n<hr>\nWeights and Biases has been used for Dataset Versioning\nLink: [WandB Artifact](https://wandb.ai/dchanda/Wikipedia/artifacts/dataset/Wiki-data/ac20734be4747897b5ba)\n![](https://i.imgur.com/3neB8pj.jpg)\n<hr>\nThe latest dataset can be downloaded using the following code:\n```\nrun = wandb.init(project=\"Wikipedia\", \n                 anonymous=\"must\")\nartifact = run.use_artifact('dchanda/Wikipedia/Wiki-data:latest', type='dataset')\nartifact_dir = artifact.download()\nrun.finish()\n\nfor file in os.listdir(artifact_dir):\n    filepath = os.path.join(artifact_dir, file)\n    with open(filepath, \"rb\") as fp:\n        contents = pickle.load(fp)\n```\n<hr>\nV0: Contains samples from file 00000-00003 from the dataset provided at  [analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/](https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/)\n\nWill add samples from more files in the corresponding versions",
      "votes": null
    },
    {
      "id": "1568842",
      "postDate": "11/03/2021 04:31:27",
      "content": "<p><strong>UPDATE</strong> 3rd Nov: Version 2 uploaded<br>\nNegatively sampled examples have been added<br>\n</p><hr><br>\n Code used to generate negatively sampled examples<p></p>\n<pre><code>negative_contents = []\n\nfor content in contents:\n    new_content = {}\n    new_content['b64_bytes'] = content['b64_bytes']\n    new_content['target'] = -1\n    c = random.choice(contents)\n    new_content['caption_title_and_reference_description'] = c['caption_title_and_reference_description']\n    if c['b64_bytes'] != content['b64_bytes']:\n        negative_contents.append(new_content)\n</code></pre>",
      "rawMarkdown": "**UPDATE** 3rd Nov: Version 2 uploaded\nNegatively sampled examples have been added\n<hr>\n Code used to generate negatively sampled examples\n```\nnegative_contents = []\n\nfor content in contents:\n    new_content = {}\n    new_content['b64_bytes'] = content['b64_bytes']\n    new_content['target'] = -1\n    c = random.choice(contents)\n    new_content['caption_title_and_reference_description'] = c['caption_title_and_reference_description']\n    if c['b64_bytes'] != content['b64_bytes']:\n        negative_contents.append(new_content)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1568842,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "11/03/2021 04:31:27",
      "content": "<p><strong>UPDATE</strong> 3rd Nov: Version 2 uploaded<br>\nNegatively sampled examples have been added<br>\n</p><hr><br>\n Code used to generate negatively sampled examples<p></p>\n<pre><code>negative_contents = []\n\nfor content in contents:\n    new_content = {}\n    new_content['b64_bytes'] = content['b64_bytes']\n    new_content['target'] = -1\n    c = random.choice(contents)\n    new_content['caption_title_and_reference_description'] = c['caption_title_and_reference_description']\n    if c['b64_bytes'] != content['b64_bytes']:\n        negative_contents.append(new_content)\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1567682": "Dataset Creation and Versioning: [https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset](https://www.kaggle.com/debarshichanda/eda-versioning-easy-to-use-dataset)\n<span>This dataset contains 300px <code>b64_bytes</code> for all the images (no downloading required) along with a <code>list of captions</code> in all the available languages</span>\n\nSome key points to notice are that not all the images have captions corresponding to them, and those samples have been filtered from this dataset.\n<hr>\nWeights and Biases has been used for Dataset Versioning\nLink: [WandB Artifact](https://wandb.ai/dchanda/Wikipedia/artifacts/dataset/Wiki-data/ac20734be4747897b5ba)\n![](https://i.imgur.com/3neB8pj.jpg)\n<hr>\nThe latest dataset can be downloaded using the following code:\n```\nrun = wandb.init(project=\"Wikipedia\", \n                 anonymous=\"must\")\nartifact = run.use_artifact('dchanda/Wikipedia/Wiki-data:latest', type='dataset')\nartifact_dir = artifact.download()\nrun.finish()\n\nfor file in os.listdir(artifact_dir):\n    filepath = os.path.join(artifact_dir, file)\n    with open(filepath, \"rb\") as fp:\n        contents = pickle.load(fp)\n```\n<hr>\nV0: Contains samples from file 00000-00003 from the dataset provided at  [analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/](https://analytics.wikimedia.org/published/datasets/one-off/caption_competition/training/joined/)\n\nWill add samples from more files in the corresponding versions",
    "1568842": "**UPDATE** 3rd Nov: Version 2 uploaded\nNegatively sampled examples have been added\n<hr>\n Code used to generate negatively sampled examples\n```\nnegative_contents = []\n\nfor content in contents:\n    new_content = {}\n    new_content['b64_bytes'] = content['b64_bytes']\n    new_content['target'] = -1\n    c = random.choice(contents)\n    new_content['caption_title_and_reference_description'] = c['caption_title_and_reference_description']\n    if c['b64_bytes'] != content['b64_bytes']:\n        negative_contents.append(new_content)\n```"
  },
  "source": "meta"
}