{
  "id": 496796,
  "title": "How to access all the test data to make a late submission?",
  "url": "/competitions/kuzushiji-recognition/discussion/496796",
  "author_name": "",
  "post_date": "2024-04-22T13:42:17.587455400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I would like to make a late submission but I don't understand how/where to access the 4150 test image files.<br>\nThe test_images.zip file publicly available in the data tab of the competition only features 1730 unlabelled test images.</p>\n<p>To test, I tried running the code below in a Kaggle notebook so see if the <code>test_images.zip</code> gets replaced by a version with the full test set when you submit a notebook as a submission. However, this approach also produces a submission file with only 1730 rows.</p>\n<pre><code>!unzip -q /kaggle//kuzushiji-recognition/test_images. -d test_images\n os\n pathlib  Path\n(, (os.listdir()))\n (, )  f:\n    f.write()\n     fname  os.listdir():\n        f.write(Path(fname).stem + )\n</code></pre>\n<p>If I had to guess, the test data is probably taken from the books highlighted in yellow in this table: <a href=\"http://codh.rois.ac.jp/char-shape/book/\" target=\"_blank\">http://codh.rois.ac.jp/char-shape/book/</a><br>\nThe notice above the table says that the books highlighted in yellow are those that were made public or updated in November 2009.<br>\nThe competition closed in October 2009.</p>\n<p>I could download that data but the file names are formatted as <code>&lt;国文研書誌ID&gt;_&lt;dual-page photo ID&gt;_&lt;1 for right page, 2 for left page&gt;.jpg</code>.<br>\nWhile Kaggle names them <code>test_&lt;long hexadecimal? string&gt;.jpg</code></p>",
  "messages": [
    {
      "id": "2767696",
      "postDate": "04/22/2024 13:42:17",
      "content": "<p>Hello,</p>\n<p>I would like to make a late submission but I don't understand how/where to access the 4150 test image files.<br>\nThe test_images.zip file publicly available in the data tab of the competition only features 1730 unlabelled test images.</p>\n<p>To test, I tried running the code below in a Kaggle notebook so see if the <code>test_images.zip</code> gets replaced by a version with the full test set when you submit a notebook as a submission. However, this approach also produces a submission file with only 1730 rows.</p>\n<pre><code>!unzip -q /kaggle//kuzushiji-recognition/test_images. -d test_images\n os\n pathlib  Path\n(, (os.listdir()))\n (, )  f:\n    f.write()\n     fname  os.listdir():\n        f.write(Path(fname).stem + )\n</code></pre>\n<p>If I had to guess, the test data is probably taken from the books highlighted in yellow in this table: <a href=\"http://codh.rois.ac.jp/char-shape/book/\" target=\"_blank\">http://codh.rois.ac.jp/char-shape/book/</a><br>\nThe notice above the table says that the books highlighted in yellow are those that were made public or updated in November 2009.<br>\nThe competition closed in October 2009.</p>\n<p>I could download that data but the file names are formatted as <code>&lt;国文研書誌ID&gt;_&lt;dual-page photo ID&gt;_&lt;1 for right page, 2 for left page&gt;.jpg</code>.<br>\nWhile Kaggle names them <code>test_&lt;long hexadecimal? string&gt;.jpg</code></p>",
      "rawMarkdown": "Hello,\n\nI would like to make a late submission but I don't understand how/where to access the 4150 test image files.\nThe test_images.zip file publicly available in the data tab of the competition only features 1730 unlabelled test images.\n\nTo test, I tried running the code below in a Kaggle notebook so see if the `test_images.zip` gets replaced by a version with the full test set when you submit a notebook as a submission. However, this approach also produces a submission file with only 1730 rows.\n\n```py\n!unzip -q /kaggle/input/kuzushiji-recognition/test_images.zip -d test_images\nimport os\nfrom pathlib import Path\nprint(\"Number of test images\", len(os.listdir(\"test_images\")))\nwith open(\"submission.csv\", \"w\") as f:\n    f.write(\"image_id,labels\\n\")\n    for fname in os.listdir(\"test_images\"):\n        f.write(Path(fname).stem + \",\\n\")\n```\n\nIf I had to guess, the test data is probably taken from the books highlighted in yellow in this table: http://codh.rois.ac.jp/char-shape/book/\nThe notice above the table says that the books highlighted in yellow are those that were made public or updated in November 2009.\nThe competition closed in October 2009.\n\nI could download that data but the file names are formatted as `<国文研書誌ID>_<dual-page photo ID>_<1 for right page, 2 for left page>.jpg`.\nWhile Kaggle names them `test_<long hexadecimal? string>.jpg`",
      "votes": null
    },
    {
      "id": "2770094",
      "postDate": "04/23/2024 17:19:50",
      "content": "<p>I wrote a Python script to find the equivalents of all the public test images, in the 日本古典籍くずし字データセット.</p>\n<pre><code> os\n multiprocessing  Pool, cpu_count\n typing  *\n tqdm  tqdm\n imageio.v3  improps, imread\n functools  reduce \n operator\n\n ():\n    n_pixels_a:  = (flat_im_r_a)\n    shape_b = improps(file_b).shape\n    n_pixels_b:  = reduce(operator.mul, shape_b[:], )\n     n_pixels_a != n_pixels_b:\n         \n    im_r_b = imread(file_b)[:, :, ]\n    thresh:  = \n    \n    \n    half_n_pixels = n_pixels_a // \n     (flat_im_r_a[half_n_pixels:] == im_r_b.flatten()[half_n_pixels:]) / half_n_pixels &gt; thresh:\n         \n\n     \n\n\n () -&gt; [[, []]]:\n    candidates = []\n    \n    \n    flat_im_r_a = imread(file_a)[:, :, ].flatten()\n     file_b  set_b:\n         compare_files(flat_im_r_a, file_b):\n            candidates.append(file_b)\n\n     (candidates) &gt; :\n         file_a, candidates\n    :\n         \n\n ():\n     (args) != :\n         ValueError()\n     find_equivalent_files(*args)\n\n () -&gt; []:\n     (( direntry: direntry.path, os.scandir()))\n\n __name__ == :\n    set_a = (listdir_with_path())\n\n    new_kokubunkenshoshi_ids = (, , , , , , , , , , , , , , , )\n\n    set_b = []\n     new_id  new_kokubunkenshoshi_ids:\n        set_b += listdir_with_path(os.path.join(, new_id, ))\n\n\n    equivalent_files = {}\n    \n     Pool(processes=cpu_count())  pool:\n         tqdm(total=(set_a), desc=)  pbar:\n             result  pool.imap_unordered(find_equivalent_files_wrapper, [(file_a, set_b)  file_a  set_a]):\n                pbar.update()\n                 result   :\n                    file_a, equivalent_b = result\n                    equivalent_files[file_a] = equivalent_b\n                    ()\n</code></pre>\n<p>When comparing the train images in the competition with the same ones from the dataset export, I noticed, to my surprise, that their pixel content ever so slightly differed, so I had to compute degrees of similarity between images of same size and keep the most similar.</p>\n<p>After post-processing, I get the mappings shown in the attached file. All test_images files are present with the exception of test_7b97f274.jpg because it's a cropped version of its associated image in the dataset export, which is 200004107_00003_2.jpg.</p>\n<p>There are pages coming from every highlighted book with the exception of The Tale of Genji.</p>\n<p>However, even if we combine all the images from all the books mentioned, we still only come up with 2040 data points. A far cry from the 4150 required…</p>",
      "rawMarkdown": "I wrote a Python script to find the equivalents of all the public test images, in the 日本古典籍くずし字データセット.\n\n```py\nimport os\nfrom multiprocessing import Pool, cpu_count\nfrom typing import *\nfrom tqdm import tqdm\nfrom imageio.v3 import improps, imread\nfrom functools import reduce # Valid in Python 2.6+, required in Python 3\nimport operator\n\ndef compare_files(flat_im_r_a, file_b):\n    n_pixels_a: int = len(flat_im_r_a)\n    shape_b = improps(file_b).shape\n    n_pixels_b: int = reduce(operator.mul, shape_b[:2], 1)\n    if n_pixels_a != n_pixels_b:\n        return False\n    im_r_b = imread(file_b)[:, :, 0]\n    thresh: float = 0.60\n    # If the half of two pictures matches sufficiently well, we can assume that the other half also will.\n    # This cuts the amount of comparisons to do by half.\n    half_n_pixels = n_pixels_a // 2\n    if sum(flat_im_r_a[half_n_pixels:] == im_r_b.flatten()[half_n_pixels:]) / half_n_pixels > thresh:\n        return True\n\n    return False\n\n\ndef find_equivalent_files(file_a, set_b) -> Optional[Tuple[str, List[str]]]:\n    candidates = []\n    # Only keep the R channel (we assume that if two pics are different, their R channels wil also greatly differ)\n    # and flatten the image (to simplify the images comparison)\n    flat_im_r_a = imread(file_a)[:, :, 0].flatten()\n    for file_b in set_b:\n        if compare_files(flat_im_r_a, file_b):\n            candidates.append(file_b)\n\n    if len(candidates) > 0:\n        return file_a, candidates\n    else:\n        return None\n\ndef find_equivalent_files_wrapper(args):\n    if len(args) != 2:\n        raise ValueError(f\"{len(args)} != 2\")\n    return find_equivalent_files(*args)\n\ndef listdir_with_path(dir: str) -> List[str]:\n    return list(map(lambda direntry: direntry.path, os.scandir(dir)))\n\nif __name__ == '__main__':\n    set_a = set(listdir_with_path(\"kuzushiji-recognition/test_images\"))\n\n    new_kokubunkenshoshi_ids = ('200003803', '200004107', '200005798', '200006665', '200008003', '200008316', '200010454', '200015843', '200017458', '200018243', '200019865', '200020019', '200021063', '200021071', '200021086', '200025191')\n\n    set_b = []\n    for new_id in new_kokubunkenshoshi_ids:\n        set_b += listdir_with_path(os.path.join(\"full-kuzushiji-dataset\", new_id, \"images\"))\n\n\n    equivalent_files = {}\n    # Use multiprocessing to speed up the process\n    with Pool(processes=cpu_count()) as pool:\n        with tqdm(total=len(set_a), desc=\"Finding equivalents\") as pbar:\n            for result in pool.imap_unordered(find_equivalent_files_wrapper, [(file_a, set_b) for file_a in set_a]):\n                pbar.update(1)\n                if result is not None:\n                    file_a, equivalent_b = result\n                    equivalent_files[file_a] = equivalent_b\n                    print(f\"{file_a} -> {equivalent_b}\")\n```\n\nWhen comparing the train images in the competition with the same ones from the dataset export, I noticed, to my surprise, that their pixel content ever so slightly differed, so I had to compute degrees of similarity between images of same size and keep the most similar.\n\nAfter post-processing, I get the mappings shown in the attached file. All test_images files are present with the exception of test_7b97f274.jpg because it's a cropped version of its associated image in the dataset export, which is 200004107_00003_2.jpg.\n\nThere are pages coming from every highlighted book with the exception of The Tale of Genji.\n\nHowever, even if we combine all the images from all the books mentioned, we still only come up with 2040 data points. A far cry from the 4150 required...",
      "votes": null
    },
    {
      "id": "2789047",
      "postDate": "05/02/2024 14:24:16",
      "content": "<p>One more clue to the mystery.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17378442%2F9fde574c3f92a85a42c7c50228905910%2Ftest_set_data_slide.png?generation=1714659828468445&amp;alt=media\"><br>\nSource: <a href=\"https://www.youtube.com/watch?v=cGpIVyV96Hg&amp;t=474\" target=\"_blank\">https://www.youtube.com/watch?v=cGpIVyV96Hg&amp;t=474</a></p>",
      "rawMarkdown": "One more clue to the mystery.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17378442%2F9fde574c3f92a85a42c7c50228905910%2Ftest_set_data_slide.png?generation=1714659828468445&alt=media)\nSource: https://www.youtube.com/watch?v=cGpIVyV96Hg&t=474",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2770094,
      "author_name": "precondition",
      "author_url": "",
      "post_date": "04/23/2024 17:19:50",
      "content": "<p>I wrote a Python script to find the equivalents of all the public test images, in the 日本古典籍くずし字データセット.</p>\n<pre><code> os\n multiprocessing  Pool, cpu_count\n typing  *\n tqdm  tqdm\n imageio.v3  improps, imread\n functools  reduce \n operator\n\n ():\n    n_pixels_a:  = (flat_im_r_a)\n    shape_b = improps(file_b).shape\n    n_pixels_b:  = reduce(operator.mul, shape_b[:], )\n     n_pixels_a != n_pixels_b:\n         \n    im_r_b = imread(file_b)[:, :, ]\n    thresh:  = \n    \n    \n    half_n_pixels = n_pixels_a // \n     (flat_im_r_a[half_n_pixels:] == im_r_b.flatten()[half_n_pixels:]) / half_n_pixels &gt; thresh:\n         \n\n     \n\n\n () -&gt; [[, []]]:\n    candidates = []\n    \n    \n    flat_im_r_a = imread(file_a)[:, :, ].flatten()\n     file_b  set_b:\n         compare_files(flat_im_r_a, file_b):\n            candidates.append(file_b)\n\n     (candidates) &gt; :\n         file_a, candidates\n    :\n         \n\n ():\n     (args) != :\n         ValueError()\n     find_equivalent_files(*args)\n\n () -&gt; []:\n     (( direntry: direntry.path, os.scandir()))\n\n __name__ == :\n    set_a = (listdir_with_path())\n\n    new_kokubunkenshoshi_ids = (, , , , , , , , , , , , , , , )\n\n    set_b = []\n     new_id  new_kokubunkenshoshi_ids:\n        set_b += listdir_with_path(os.path.join(, new_id, ))\n\n\n    equivalent_files = {}\n    \n     Pool(processes=cpu_count())  pool:\n         tqdm(total=(set_a), desc=)  pbar:\n             result  pool.imap_unordered(find_equivalent_files_wrapper, [(file_a, set_b)  file_a  set_a]):\n                pbar.update()\n                 result   :\n                    file_a, equivalent_b = result\n                    equivalent_files[file_a] = equivalent_b\n                    ()\n</code></pre>\n<p>When comparing the train images in the competition with the same ones from the dataset export, I noticed, to my surprise, that their pixel content ever so slightly differed, so I had to compute degrees of similarity between images of same size and keep the most similar.</p>\n<p>After post-processing, I get the mappings shown in the attached file. All test_images files are present with the exception of test_7b97f274.jpg because it's a cropped version of its associated image in the dataset export, which is 200004107_00003_2.jpg.</p>\n<p>There are pages coming from every highlighted book with the exception of The Tale of Genji.</p>\n<p>However, even if we combine all the images from all the books mentioned, we still only come up with 2040 data points. A far cry from the 4150 required…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2789047,
      "author_name": "precondition",
      "author_url": "",
      "post_date": "05/02/2024 14:24:16",
      "content": "<p>One more clue to the mystery.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17378442%2F9fde574c3f92a85a42c7c50228905910%2Ftest_set_data_slide.png?generation=1714659828468445&amp;alt=media\"><br>\nSource: <a href=\"https://www.youtube.com/watch?v=cGpIVyV96Hg&amp;t=474\" target=\"_blank\">https://www.youtube.com/watch?v=cGpIVyV96Hg&amp;t=474</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2767696": "Hello,\n\nI would like to make a late submission but I don't understand how/where to access the 4150 test image files.\nThe test_images.zip file publicly available in the data tab of the competition only features 1730 unlabelled test images.\n\nTo test, I tried running the code below in a Kaggle notebook so see if the `test_images.zip` gets replaced by a version with the full test set when you submit a notebook as a submission. However, this approach also produces a submission file with only 1730 rows.\n\n```py\n!unzip -q /kaggle/input/kuzushiji-recognition/test_images.zip -d test_images\nimport os\nfrom pathlib import Path\nprint(\"Number of test images\", len(os.listdir(\"test_images\")))\nwith open(\"submission.csv\", \"w\") as f:\n    f.write(\"image_id,labels\\n\")\n    for fname in os.listdir(\"test_images\"):\n        f.write(Path(fname).stem + \",\\n\")\n```\n\nIf I had to guess, the test data is probably taken from the books highlighted in yellow in this table: http://codh.rois.ac.jp/char-shape/book/\nThe notice above the table says that the books highlighted in yellow are those that were made public or updated in November 2009.\nThe competition closed in October 2009.\n\nI could download that data but the file names are formatted as `<国文研書誌ID>_<dual-page photo ID>_<1 for right page, 2 for left page>.jpg`.\nWhile Kaggle names them `test_<long hexadecimal? string>.jpg`",
    "2770094": "I wrote a Python script to find the equivalents of all the public test images, in the 日本古典籍くずし字データセット.\n\n```py\nimport os\nfrom multiprocessing import Pool, cpu_count\nfrom typing import *\nfrom tqdm import tqdm\nfrom imageio.v3 import improps, imread\nfrom functools import reduce # Valid in Python 2.6+, required in Python 3\nimport operator\n\ndef compare_files(flat_im_r_a, file_b):\n    n_pixels_a: int = len(flat_im_r_a)\n    shape_b = improps(file_b).shape\n    n_pixels_b: int = reduce(operator.mul, shape_b[:2], 1)\n    if n_pixels_a != n_pixels_b:\n        return False\n    im_r_b = imread(file_b)[:, :, 0]\n    thresh: float = 0.60\n    # If the half of two pictures matches sufficiently well, we can assume that the other half also will.\n    # This cuts the amount of comparisons to do by half.\n    half_n_pixels = n_pixels_a // 2\n    if sum(flat_im_r_a[half_n_pixels:] == im_r_b.flatten()[half_n_pixels:]) / half_n_pixels > thresh:\n        return True\n\n    return False\n\n\ndef find_equivalent_files(file_a, set_b) -> Optional[Tuple[str, List[str]]]:\n    candidates = []\n    # Only keep the R channel (we assume that if two pics are different, their R channels wil also greatly differ)\n    # and flatten the image (to simplify the images comparison)\n    flat_im_r_a = imread(file_a)[:, :, 0].flatten()\n    for file_b in set_b:\n        if compare_files(flat_im_r_a, file_b):\n            candidates.append(file_b)\n\n    if len(candidates) > 0:\n        return file_a, candidates\n    else:\n        return None\n\ndef find_equivalent_files_wrapper(args):\n    if len(args) != 2:\n        raise ValueError(f\"{len(args)} != 2\")\n    return find_equivalent_files(*args)\n\ndef listdir_with_path(dir: str) -> List[str]:\n    return list(map(lambda direntry: direntry.path, os.scandir(dir)))\n\nif __name__ == '__main__':\n    set_a = set(listdir_with_path(\"kuzushiji-recognition/test_images\"))\n\n    new_kokubunkenshoshi_ids = ('200003803', '200004107', '200005798', '200006665', '200008003', '200008316', '200010454', '200015843', '200017458', '200018243', '200019865', '200020019', '200021063', '200021071', '200021086', '200025191')\n\n    set_b = []\n    for new_id in new_kokubunkenshoshi_ids:\n        set_b += listdir_with_path(os.path.join(\"full-kuzushiji-dataset\", new_id, \"images\"))\n\n\n    equivalent_files = {}\n    # Use multiprocessing to speed up the process\n    with Pool(processes=cpu_count()) as pool:\n        with tqdm(total=len(set_a), desc=\"Finding equivalents\") as pbar:\n            for result in pool.imap_unordered(find_equivalent_files_wrapper, [(file_a, set_b) for file_a in set_a]):\n                pbar.update(1)\n                if result is not None:\n                    file_a, equivalent_b = result\n                    equivalent_files[file_a] = equivalent_b\n                    print(f\"{file_a} -> {equivalent_b}\")\n```\n\nWhen comparing the train images in the competition with the same ones from the dataset export, I noticed, to my surprise, that their pixel content ever so slightly differed, so I had to compute degrees of similarity between images of same size and keep the most similar.\n\nAfter post-processing, I get the mappings shown in the attached file. All test_images files are present with the exception of test_7b97f274.jpg because it's a cropped version of its associated image in the dataset export, which is 200004107_00003_2.jpg.\n\nThere are pages coming from every highlighted book with the exception of The Tale of Genji.\n\nHowever, even if we combine all the images from all the books mentioned, we still only come up with 2040 data points. A far cry from the 4150 required...",
    "2789047": "One more clue to the mystery.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17378442%2F9fde574c3f92a85a42c7c50228905910%2Ftest_set_data_slide.png?generation=1714659828468445&alt=media)\nSource: https://www.youtube.com/watch?v=cGpIVyV96Hg&t=474"
  },
  "source": "meta"
}