{
  "id": 30749,
  "title": "Announcement: Additional data updated",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30749",
  "author_name": "",
  "post_date": "2017-03-27T23:12:45.825666100Z",
  "votes": 7,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>As described in <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30621\">this</a> thread, we had fixed the labeling errors in the additional dataset. </p>\n\n<ol>\n<li><p>If you had downloaded additional.7z file before March 24, 2017, you can use two csv files to avoid re-downloading. fixed_labels has the file names, their old labels and new labels. removed_files has the list of files to remove. </p></li>\n<li><p>If you haven't downloaded the additional dataset, they are now available (with the correct labels and no duplicates) as 3 parts: additional_Type_{x}.7z, x=1,2,3. </p></li>\n<li><p>We don't officially support the torrent download for this competition, so I have pulled down the torrent file. If someone in the community wants to host, please send me the torrent file and I'll upload.</p></li>\n<li><p>Both the Intel server and Kaggle Kernels will have the updated additional data soon. </p></li>\n</ol>\n\n<p>Thank you all for your patience. </p>\n\n<p>Kaggle admin</p>",
  "messages": [
    {
      "id": "170904",
      "postDate": "03/27/2017 23:12:45",
      "content": "<p>Hi all,</p>\n\n<p>As described in <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30621\">this</a> thread, we had fixed the labeling errors in the additional dataset. </p>\n\n<ol>\n<li><p>If you had downloaded additional.7z file before March 24, 2017, you can use two csv files to avoid re-downloading. fixed_labels has the file names, their old labels and new labels. removed_files has the list of files to remove. </p></li>\n<li><p>If you haven't downloaded the additional dataset, they are now available (with the correct labels and no duplicates) as 3 parts: additional_Type_{x}.7z, x=1,2,3. </p></li>\n<li><p>We don't officially support the torrent download for this competition, so I have pulled down the torrent file. If someone in the community wants to host, please send me the torrent file and I'll upload.</p></li>\n<li><p>Both the Intel server and Kaggle Kernels will have the updated additional data soon. </p></li>\n</ol>\n\n<p>Thank you all for your patience. </p>\n\n<p>Kaggle admin</p>",
      "rawMarkdown": "Hi all,\n\nAs described in [this][1] thread, we had fixed the labeling errors in the additional dataset. \n\n1. If you had downloaded additional.7z file before March 24, 2017, you can use two csv files to avoid re-downloading. fixed_labels has the file names, their old labels and new labels. removed_files has the list of files to remove. \n\n2. If you haven't downloaded the additional dataset, they are now available (with the correct labels and no duplicates) as 3 parts: additional_Type_{x}.7z, x=1,2,3. \n\n3. We don't officially support the torrent download for this competition, so I have pulled down the torrent file. If someone in the community wants to host, please send me the torrent file and I'll upload.\n\n4. Both the Intel server and Kaggle Kernels will have the updated additional data soon. \n\nThank you all for your patience. \n\nKaggle admin\n\n\n  [1]: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30621",
      "votes": null
    },
    {
      "id": "170908",
      "postDate": "03/27/2017 23:44:04",
      "content": "<p>How to wget those datasets on an ec2 instance? I tried:</p>\n\n<p><code>wget https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/download/train.7z</code></p>\n\n<p>but it doesn't download actual dataset</p>",
      "rawMarkdown": "How to wget those datasets on an ec2 instance? I tried:\n\n`wget https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/download/train.7z`\n\nbut it doesn't download actual dataset",
      "votes": null
    },
    {
      "id": "170933",
      "postDate": "03/28/2017 02:13:26",
      "content": "<p>You may try this:</p>\n\n<p><code>\nwget https://kaggle2.blob.core.windows.net/competitions-data/kaggle/6243/train.7z?sv=2015-12-11&amp;sr=b&amp;sig=LuJRoYHig6Df88fnn5SaXFVO6RjT8O6TswVfD%2FrE46M%3D&amp;se=2017-03-31T02%3A12%3A39Z&amp;sp=r\n</code></p>",
      "rawMarkdown": "You may try this:\n\n```\nwget https://kaggle2.blob.core.windows.net/competitions-data/kaggle/6243/train.7z?sv=2015-12-11&sr=b&sig=LuJRoYHig6Df88fnn5SaXFVO6RjT8O6TswVfD%2FrE46M%3D&se=2017-03-31T02%3A12%3A39Z&sp=r\n```",
      "votes": null
    },
    {
      "id": "170946",
      "postDate": "03/28/2017 03:22:43",
      "content": "<p>It gives me 404 error</p>",
      "rawMarkdown": "It gives me 404 error",
      "votes": null
    },
    {
      "id": "170981",
      "postDate": "03/28/2017 07:45:39",
      "content": "<p>thank you</p>",
      "rawMarkdown": "thank you",
      "votes": null
    },
    {
      "id": "170983",
      "postDate": "03/28/2017 08:03:32",
      "content": "<p>Thanks so much! 😀 <br>\nI want to announce once more, when you have done updating datasets (#4.).</p>",
      "rawMarkdown": "Thanks so much! 😀  \nI want to announce once more, when you have done updating datasets (#4.).",
      "votes": null
    },
    {
      "id": "171099",
      "postDate": "03/28/2017 16:54:13",
      "content": "<p>you can use kaggle-cli\n<a href=\"https://github.com/floydwch/kaggle-cli\">https://github.com/floydwch/kaggle-cli</a></p>",
      "rawMarkdown": "you can use kaggle-cli\nhttps://github.com/floydwch/kaggle-cli",
      "votes": null
    },
    {
      "id": "171133",
      "postDate": "03/28/2017 19:31:58",
      "content": "<p>I tried to extract 7z files on Colfax Cluster, but additional_Type_2.7z are broken.</p>\n\n<p>I'm downloading at <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/data\">web page of kaggle</a>. <br>\nI hope web page's file is not broken:)</p>\n\n<hr>\n\n<p>$ md5sum /data/kaggle_3.27/additional_Type_2.7z</p>\n\n<p>38bf446a2b66a2d20fc7a4c6d9089131 /data/kaggle_3.27/additional_Type_2.7z</p>\n\n<p>$ 7za x /data/kaggle_3.27/additional_Type_2.7z</p>\n\n<p>7-Zip (a) [64] 15.09 beta : Copyright (c) 1999-2015 Igor Pavlov : 2015-10-16\np7zip Version 15.09 beta (locale=ja_JP.UTF-8,Utf16=on,HugeFiles=on,64 bits,8 CPUs Intel Core Processor (Broadwell) (306D2),ASM,AES-NI)</p>\n\n<p>Scanning the drive for archives:\n1 file, 18491891712 bytes (18 GiB)</p>\n\n<p>Extracting archive: /data/kaggle_3.27/additional_Type_2.7z\nERROR: /data/kaggle_3.27/additional_Type_2.7z\n/data/kaggle_3.27/additional_Type_2.7z\nOpen ERROR: Can not open the file as [7z] archive</p>\n\n<p>ERRORS:\nHeaders Error\nWARNINGS:\nThere are data after the end of archive</p>\n\n<p>Can't open as archive: 1\nFiles: 0\nSize:       0\nCompressed: 0</p>",
      "rawMarkdown": "I tried to extract 7z files on Colfax Cluster, but additional_Type_2.7z are broken.\n\nI'm downloading at [web page of kaggle](https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/data).  \nI hope web page's file is not broken:)\n\n----------\n$ md5sum /data/kaggle_3.27/additional_Type_2.7z\n\n38bf446a2b66a2d20fc7a4c6d9089131 /data/kaggle_3.27/additional_Type_2.7z\n\n$ 7za x /data/kaggle_3.27/additional_Type_2.7z\n\n7-Zip (a) [64] 15.09 beta : Copyright (c) 1999-2015 Igor Pavlov : 2015-10-16\np7zip Version 15.09 beta (locale=ja_JP.UTF-8,Utf16=on,HugeFiles=on,64 bits,8 CPUs Intel Core Processor (Broadwell) (306D2),ASM,AES-NI)\n\nScanning the drive for archives:\n1 file, 18491891712 bytes (18 GiB)\n\nExtracting archive: /data/kaggle_3.27/additional_Type_2.7z\nERROR: /data/kaggle_3.27/additional_Type_2.7z\n/data/kaggle_3.27/additional_Type_2.7z\nOpen ERROR: Can not open the file as [7z] archive\n\n\nERRORS:\nHeaders Error\nWARNINGS:\nThere are data after the end of archive\n    \nCan't open as archive: 1\nFiles: 0\nSize:       0\nCompressed: 0",
      "votes": null
    },
    {
      "id": "171141",
      "postDate": "03/28/2017 19:49:05",
      "content": "<p>Cool! thank you that works!</p>",
      "rawMarkdown": "Cool! thank you that works!",
      "votes": null
    },
    {
      "id": "171175",
      "postDate": "03/28/2017 23:11:07",
      "content": "<p>I have extracted additional_Type_2.7z without error, downloaded on kaggle's web page. <br>\nSorry for irregular way.</p>\n\n<ul>\n<li><p>on kaggle's web page <br>\nfilesize = 15871727173 <br>\nmd5sum = ee6c378c086bb3e77cda29b1bf4828d6 <br>\nsha1sum = 4495bf9e7de477ba64daf01e3be6e04ecbc6151a</p></li>\n<li><p>on Colfax Cluster (broken) <br>\nfilesize = 18491891712 <br>\nmd5sum = 38bf446a2b66a2d20fc7a4c6d9089131 <br>\nsha1sum = 1721366bea847c2dcd1a84cbe84453b51695daad </p></li>\n</ul>",
      "rawMarkdown": "I have extracted additional_Type_2.7z without error, downloaded on kaggle's web page.  \nSorry for irregular way.\n\n* on kaggle's web page  \n  filesize = 15871727173  \n  md5sum = ee6c378c086bb3e77cda29b1bf4828d6  \n  sha1sum = 4495bf9e7de477ba64daf01e3be6e04ecbc6151a\n\n* on Colfax Cluster (broken)  \n  filesize = 18491891712  \n  md5sum = 38bf446a2b66a2d20fc7a4c6d9089131  \n  sha1sum = 1721366bea847c2dcd1a84cbe84453b51695daad",
      "votes": null
    },
    {
      "id": "171184",
      "postDate": "03/29/2017 00:26:44",
      "content": "<p>Thanks for point out the additional_Type_2.7z with broken checksum on Colfax. We have fixed it now.</p>",
      "rawMarkdown": "Thanks for point out the additional_Type_2.7z with broken checksum on Colfax. We have fixed it now.",
      "votes": null
    },
    {
      "id": "171254",
      "postDate": "03/29/2017 07:24:13",
      "content": "<p>Sorry to bother you again. <br>\n<strong>I found duplication on removed_files.csv and fixed_labels.csv files.</strong>  </p>\n\n<p>Duplicate Remove <br>\n./Type_3/6821.jpg  </p>\n\n<p>Remove or Fix? <br>\n./Type_2/1018.jpg, ./Type_2/1094.jpg, ./Type_2/114.jpg,  ./Type_3/1295.jpg, ./Type_2/1319.jpg, <br>\n./Type_2/1510.jpg, ./Type_3/1762.jpg, ./Type_2/1799.jpg, ./Type_3/1822.jpg, ./Type_1/1825.jpg, <br>\n./Type_2/2010.jpg, ./Type_1/2151.jpg, ./Type_1/2320.jpg, ./Type_1/2414.jpg, ./Type_3/2491.jpg, <br>\n./Type_3/2511.jpg, ./Type_2/2562.jpg, ./Type_3/2613.jpg, ./Type_2/2663.jpg, ./Type_2/2860.jpg, <br>\n./Type_2/2937.jpg, ./Type_3/2998.jpg, ./Type_2/3017.jpg, ./Type_3/304.jpg,  ./Type_1/3355.jpg, <br>\n./Type_2/3615.jpg, ./Type_1/3662.jpg, ./Type_3/3692.jpg, ./Type_3/3849.jpg, ./Type_3/3856.jpg, <br>\n./Type_1/3897.jpg, ./Type_3/423.jpg,  ./Type_2/474.jpg,  ./Type_2/502.jpg,  ./Type_2/664.jpg, <br>\n./Type_2/68.jpg,   ./Type_2/811.jpg,  ./Type_3/821.jpg,  ./Type_2/989.jpg  </p>\n\n<p>So, I check the number of additional images <br>\n(removed_and_fixed means removed -&gt; fixed, removed_and_fixed means fixed -&gt; removed). <br>\nI wrote the rough script (attachment) on python 2.7. <br>\n<strong>This script is destructive. <br>\nPlease backup original additional directory!</strong>  </p>\n\n<p>name,Type_1,Type_2,Type_3,total <br>\nadditional_old,   1192,3619,2113,6924 <br>\nkaggle_kernel_ago,1192,3619,2113,6924 <br>\ncolfax_old,       1192,3619,2113,6924 <br>\nre-download,      1125,3617,2030,6772 <br>\nkaggle_kernel_now,1125,3617,2030,6772 <br>\ncolfax_now,       NaN, NaN, NaN, NaN <br>\nremoved_and_fixed,1189,3565,1974,6728 <br>\nfixed_and_removed,1201,3580,1986,6767  </p>\n\n<p>I can't check /data/kaggle_3.27/Type_* on Colfax Cluster, because of permission denied.  </p>\n\n<p>/data/kaggle_3.27 <br>\ndrwxr-x---. 2 root root 32768  3月 27 14:35 Type_1 <br>\ndrwxr-x---. 2 root root 73728  3月 27 14:35 Type_2 <br>\ndrwxr-x---. 2 root root 49152  3月 27 14:35 Type_3  </p>\n\n<p>/data/kaggle/additional <br>\ndrwxr-xr-x. 2 root root 36864  3月  8 15:15 Type_1 <br>\ndrwxr-xr-x. 2 root root 77824  3月  8 15:15 Type_2 <br>\ndrwxr-xr-x. 2 root root 53248  3月  8 15:15 Type_3  </p>\n\n<p><strong>I think removed_files.csv and fixed_labels.csv are wrong, because of duplication.</strong> <br>\nI think it is good to use re-download dataset. <br>\n<strong>Any time suits, please announce us which is the correct dataset:)</strong>  </p>\n\n<p>ADD: I do not mean to be rude, but I wrote the script (attachment) to make csv files (attachment).  </p>",
      "rawMarkdown": "Sorry to bother you again.  \n**I found duplication on removed_files.csv and fixed_labels.csv files.**  \n\nDuplicate Remove  \n./Type_3/6821.jpg  \n\nRemove or Fix?  \n./Type_2/1018.jpg, ./Type_2/1094.jpg, ./Type_2/114.jpg,  ./Type_3/1295.jpg, ./Type_2/1319.jpg,  \n./Type_2/1510.jpg, ./Type_3/1762.jpg, ./Type_2/1799.jpg, ./Type_3/1822.jpg, ./Type_1/1825.jpg,  \n./Type_2/2010.jpg, ./Type_1/2151.jpg, ./Type_1/2320.jpg, ./Type_1/2414.jpg, ./Type_3/2491.jpg,  \n./Type_3/2511.jpg, ./Type_2/2562.jpg, ./Type_3/2613.jpg, ./Type_2/2663.jpg, ./Type_2/2860.jpg,  \n./Type_2/2937.jpg, ./Type_3/2998.jpg, ./Type_2/3017.jpg, ./Type_3/304.jpg,  ./Type_1/3355.jpg,  \n./Type_2/3615.jpg, ./Type_1/3662.jpg, ./Type_3/3692.jpg, ./Type_3/3849.jpg, ./Type_3/3856.jpg,  \n./Type_1/3897.jpg, ./Type_3/423.jpg,  ./Type_2/474.jpg,  ./Type_2/502.jpg,  ./Type_2/664.jpg,  \n./Type_2/68.jpg,   ./Type_2/811.jpg,  ./Type_3/821.jpg,  ./Type_2/989.jpg  \n\nSo, I check the number of additional images  \n(removed_and_fixed means removed -> fixed, removed_and_fixed means fixed -> removed).  \nI wrote the rough script (attachment) on python 2.7.  \n**This script is destructive.  \nPlease backup original additional directory!**  \n\nname,Type_1,Type_2,Type_3,total  \nadditional_old,   1192,3619,2113,6924  \nkaggle_kernel_ago,1192,3619,2113,6924  \ncolfax_old,       1192,3619,2113,6924  \nre-download,      1125,3617,2030,6772  \nkaggle_kernel_now,1125,3617,2030,6772  \ncolfax_now,       NaN, NaN, NaN, NaN  \nremoved_and_fixed,1189,3565,1974,6728  \nfixed_and_removed,1201,3580,1986,6767  \n\nI can't check /data/kaggle_3.27/Type_* on Colfax Cluster, because of permission denied.  \n\n/data/kaggle_3.27  \ndrwxr-x---. 2 root root 32768  3月 27 14:35 Type_1  \ndrwxr-x---. 2 root root 73728  3月 27 14:35 Type_2  \ndrwxr-x---. 2 root root 49152  3月 27 14:35 Type_3  \n\n/data/kaggle/additional  \ndrwxr-xr-x. 2 root root 36864  3月  8 15:15 Type_1  \ndrwxr-xr-x. 2 root root 77824  3月  8 15:15 Type_2  \ndrwxr-xr-x. 2 root root 53248  3月  8 15:15 Type_3  \n\n**I think removed_files.csv and fixed_labels.csv are wrong, because of duplication.**  \nI think it is good to use re-download dataset.  \n**Any time suits, please announce us which is the correct dataset:)**  \n\nADD: I do not mean to be rude, but I wrote the script (attachment) to make csv files (attachment).",
      "votes": null
    },
    {
      "id": "171284",
      "postDate": "03/29/2017 09:40:27",
      "content": "<p>has the data been updated on kaggle kernels?</p>",
      "rawMarkdown": "has the data been updated on kaggle kernels?",
      "votes": null
    },
    {
      "id": "171664",
      "postDate": "03/30/2017 23:30:40",
      "content": "<p>I think there is something wrong in these new additional data and fixed label file. </p>\n\n<p>For example, the fixed_label.csv tells me 10.jpg has to be type_1, but it is still in additional_Type_2.</p>\n\n<p>Also, as written by Kambarakun, there are some duplication in file names.</p>\n\n<p>It should be fixed, I think.</p>",
      "rawMarkdown": "I think there is something wrong in these new additional data and fixed label file. \n\nFor example, the fixed_label.csv tells me 10.jpg has to be type_1, but it is still in additional_Type_2.\n\nAlso, as written by Kambarakun, there are some duplication in file names.\n\nIt should be fixed, I think.",
      "votes": null
    },
    {
      "id": "171687",
      "postDate": "03/31/2017 02:42:51",
      "content": "<p>Thanks for spotting this! This is fixed now. Files are now postfixed with <code>_v2</code>. </p>",
      "rawMarkdown": "Thanks for spotting this! This is fixed now. Files are now postfixed with `_v2`.",
      "votes": null
    },
    {
      "id": "171688",
      "postDate": "03/31/2017 02:43:02",
      "content": "<p>Thanks for spotting this! This is fixed now. Files are now postfixed with <code>_v2</code>.</p>",
      "rawMarkdown": "Thanks for spotting this! This is fixed now. Files are now postfixed with `_v2`.",
      "votes": null
    },
    {
      "id": "171695",
      "postDate": "03/31/2017 03:18:02",
      "content": "<p>Thank you so much for quick correction~!</p>",
      "rawMarkdown": "Thank you so much for quick correction~!",
      "votes": null
    },
    {
      "id": "172306",
      "postDate": "04/03/2017 07:06:05",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks.",
      "votes": null
    },
    {
      "id": "172376",
      "postDate": "04/03/2017 13:57:14",
      "content": "<p>I think there is still a discrepancy in the data. I've downloaded the <strong>train</strong> data/images and the fixed_labels_v2 on 2nd April 2017. For example 1001.jpg is still available in the 'Type_2' folder but the new label should be Type_1 , or do some images have the same id in train and additional data?</p>",
      "rawMarkdown": "I think there is still a discrepancy in the data. I've downloaded the **train** data/images and the fixed_labels_v2 on 2nd April 2017. For example 1001.jpg is still available in the 'Type_2' folder but the new label should be Type_1 , or do some images have the same id in train and additional data?",
      "votes": null
    },
    {
      "id": "172472",
      "postDate": "04/03/2017 20:51:32",
      "content": "<p>Correct - images have the same ids in train/test and additional since they all start from 0.jpg. All the fixed_labels are on the additional dataset, not train. </p>",
      "rawMarkdown": "Correct - images have the same ids in train/test and additional since they all start from 0.jpg. All the fixed_labels are on the additional dataset, not train.",
      "votes": null
    },
    {
      "id": "173836",
      "postDate": "04/08/2017 22:38:46",
      "content": "<p>Can we please get a torrent somehow? I've downloaded the data twice and I'm still getting corrupted files.</p>",
      "rawMarkdown": "Can we please get a torrent somehow? I've downloaded the data twice and I'm still getting corrupted files.",
      "votes": null
    },
    {
      "id": "174248",
      "postDate": "04/10/2017 19:56:09",
      "content": "<p>Hello anybody wanna try out the torrent file I included?\nI have never tried to create a torrent before.</p>\n\n<p>If this is working, I'll try to create more(or a single one) torrent after I succeed in downloading all the big zip files tomorrow.</p>",
      "rawMarkdown": "Hello anybody wanna try out the torrent file I included?\nI have never tried to create a torrent before.\n\nIf this is working, I'll try to create more(or a single one) torrent after I succeed in downloading all the big zip files tomorrow.",
      "votes": null
    },
    {
      "id": "175144",
      "postDate": "04/14/2017 01:19:38",
      "content": "<p>Hi everybody, I've make these 5 big image zip files into a torrent.\nI am not sure if hosting this will block my internet or something.\nAnyone who hasn't successfully downloaded could give it a try.</p>",
      "rawMarkdown": "Hi everybody, I've make these 5 big image zip files into a torrent.\nI am not sure if hosting this will block my internet or something.\nAnyone who hasn't successfully downloaded could give it a try.",
      "votes": null
    },
    {
      "id": "188554",
      "postDate": "06/03/2017 07:24:03",
      "content": "<p>Hi, Does the torrent file still work? It seems to fail.</p>",
      "rawMarkdown": "Hi, Does the torrent file still work? It seems to fail.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 170908,
      "author_name": "herimanitra",
      "author_url": "",
      "post_date": "03/27/2017 23:44:04",
      "content": "<p>How to wget those datasets on an ec2 instance? I tried:</p>\n\n<p><code>wget https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/download/train.7z</code></p>\n\n<p>but it doesn't download actual dataset</p>",
      "votes": null,
      "replies": [
        {
          "id": 170933,
          "author_name": "duinodu",
          "author_url": "",
          "post_date": "03/28/2017 02:13:26",
          "content": "<p>You may try this:</p>\n\n<p><code>\nwget https://kaggle2.blob.core.windows.net/competitions-data/kaggle/6243/train.7z?sv=2015-12-11&amp;sr=b&amp;sig=LuJRoYHig6Df88fnn5SaXFVO6RjT8O6TswVfD%2FrE46M%3D&amp;se=2017-03-31T02%3A12%3A39Z&amp;sp=r\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 170946,
          "author_name": "herimanitra",
          "author_url": "",
          "post_date": "03/28/2017 03:22:43",
          "content": "<p>It gives me 404 error</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 171099,
          "author_name": "ronigur",
          "author_url": "",
          "post_date": "03/28/2017 16:54:13",
          "content": "<p>you can use kaggle-cli\n<a href=\"https://github.com/floydwch/kaggle-cli\">https://github.com/floydwch/kaggle-cli</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 171141,
          "author_name": "herimanitra",
          "author_url": "",
          "post_date": "03/28/2017 19:49:05",
          "content": "<p>Cool! thank you that works!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 170981,
      "author_name": "steelrose",
      "author_url": "",
      "post_date": "03/28/2017 07:45:39",
      "content": "<p>thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 170983,
      "author_name": "kambarakun",
      "author_url": "",
      "post_date": "03/28/2017 08:03:32",
      "content": "<p>Thanks so much! 😀 <br>\nI want to announce once more, when you have done updating datasets (#4.).</p>",
      "votes": null,
      "replies": [
        {
          "id": 171133,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "03/28/2017 19:31:58",
          "content": "<p>I tried to extract 7z files on Colfax Cluster, but additional_Type_2.7z are broken.</p>\n\n<p>I'm downloading at <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/data\">web page of kaggle</a>. <br>\nI hope web page's file is not broken:)</p>\n\n<hr>\n\n<p>$ md5sum /data/kaggle_3.27/additional_Type_2.7z</p>\n\n<p>38bf446a2b66a2d20fc7a4c6d9089131 /data/kaggle_3.27/additional_Type_2.7z</p>\n\n<p>$ 7za x /data/kaggle_3.27/additional_Type_2.7z</p>\n\n<p>7-Zip (a) [64] 15.09 beta : Copyright (c) 1999-2015 Igor Pavlov : 2015-10-16\np7zip Version 15.09 beta (locale=ja_JP.UTF-8,Utf16=on,HugeFiles=on,64 bits,8 CPUs Intel Core Processor (Broadwell) (306D2),ASM,AES-NI)</p>\n\n<p>Scanning the drive for archives:\n1 file, 18491891712 bytes (18 GiB)</p>\n\n<p>Extracting archive: /data/kaggle_3.27/additional_Type_2.7z\nERROR: /data/kaggle_3.27/additional_Type_2.7z\n/data/kaggle_3.27/additional_Type_2.7z\nOpen ERROR: Can not open the file as [7z] archive</p>\n\n<p>ERRORS:\nHeaders Error\nWARNINGS:\nThere are data after the end of archive</p>\n\n<p>Can't open as archive: 1\nFiles: 0\nSize:       0\nCompressed: 0</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 171175,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "03/28/2017 23:11:07",
          "content": "<p>I have extracted additional_Type_2.7z without error, downloaded on kaggle's web page. <br>\nSorry for irregular way.</p>\n\n<ul>\n<li><p>on kaggle's web page <br>\nfilesize = 15871727173 <br>\nmd5sum = ee6c378c086bb3e77cda29b1bf4828d6 <br>\nsha1sum = 4495bf9e7de477ba64daf01e3be6e04ecbc6151a</p></li>\n<li><p>on Colfax Cluster (broken) <br>\nfilesize = 18491891712 <br>\nmd5sum = 38bf446a2b66a2d20fc7a4c6d9089131 <br>\nsha1sum = 1721366bea847c2dcd1a84cbe84453b51695daad </p></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 171184,
          "author_name": "kumarhemanth",
          "author_url": "",
          "post_date": "03/29/2017 00:26:44",
          "content": "<p>Thanks for point out the additional_Type_2.7z with broken checksum on Colfax. We have fixed it now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 171254,
      "author_name": "kambarakun",
      "author_url": "",
      "post_date": "03/29/2017 07:24:13",
      "content": "<p>Sorry to bother you again. <br>\n<strong>I found duplication on removed_files.csv and fixed_labels.csv files.</strong>  </p>\n\n<p>Duplicate Remove <br>\n./Type_3/6821.jpg  </p>\n\n<p>Remove or Fix? <br>\n./Type_2/1018.jpg, ./Type_2/1094.jpg, ./Type_2/114.jpg,  ./Type_3/1295.jpg, ./Type_2/1319.jpg, <br>\n./Type_2/1510.jpg, ./Type_3/1762.jpg, ./Type_2/1799.jpg, ./Type_3/1822.jpg, ./Type_1/1825.jpg, <br>\n./Type_2/2010.jpg, ./Type_1/2151.jpg, ./Type_1/2320.jpg, ./Type_1/2414.jpg, ./Type_3/2491.jpg, <br>\n./Type_3/2511.jpg, ./Type_2/2562.jpg, ./Type_3/2613.jpg, ./Type_2/2663.jpg, ./Type_2/2860.jpg, <br>\n./Type_2/2937.jpg, ./Type_3/2998.jpg, ./Type_2/3017.jpg, ./Type_3/304.jpg,  ./Type_1/3355.jpg, <br>\n./Type_2/3615.jpg, ./Type_1/3662.jpg, ./Type_3/3692.jpg, ./Type_3/3849.jpg, ./Type_3/3856.jpg, <br>\n./Type_1/3897.jpg, ./Type_3/423.jpg,  ./Type_2/474.jpg,  ./Type_2/502.jpg,  ./Type_2/664.jpg, <br>\n./Type_2/68.jpg,   ./Type_2/811.jpg,  ./Type_3/821.jpg,  ./Type_2/989.jpg  </p>\n\n<p>So, I check the number of additional images <br>\n(removed_and_fixed means removed -&gt; fixed, removed_and_fixed means fixed -&gt; removed). <br>\nI wrote the rough script (attachment) on python 2.7. <br>\n<strong>This script is destructive. <br>\nPlease backup original additional directory!</strong>  </p>\n\n<p>name,Type_1,Type_2,Type_3,total <br>\nadditional_old,   1192,3619,2113,6924 <br>\nkaggle_kernel_ago,1192,3619,2113,6924 <br>\ncolfax_old,       1192,3619,2113,6924 <br>\nre-download,      1125,3617,2030,6772 <br>\nkaggle_kernel_now,1125,3617,2030,6772 <br>\ncolfax_now,       NaN, NaN, NaN, NaN <br>\nremoved_and_fixed,1189,3565,1974,6728 <br>\nfixed_and_removed,1201,3580,1986,6767  </p>\n\n<p>I can't check /data/kaggle_3.27/Type_* on Colfax Cluster, because of permission denied.  </p>\n\n<p>/data/kaggle_3.27 <br>\ndrwxr-x---. 2 root root 32768  3月 27 14:35 Type_1 <br>\ndrwxr-x---. 2 root root 73728  3月 27 14:35 Type_2 <br>\ndrwxr-x---. 2 root root 49152  3月 27 14:35 Type_3  </p>\n\n<p>/data/kaggle/additional <br>\ndrwxr-xr-x. 2 root root 36864  3月  8 15:15 Type_1 <br>\ndrwxr-xr-x. 2 root root 77824  3月  8 15:15 Type_2 <br>\ndrwxr-xr-x. 2 root root 53248  3月  8 15:15 Type_3  </p>\n\n<p><strong>I think removed_files.csv and fixed_labels.csv are wrong, because of duplication.</strong> <br>\nI think it is good to use re-download dataset. <br>\n<strong>Any time suits, please announce us which is the correct dataset:)</strong>  </p>\n\n<p>ADD: I do not mean to be rude, but I wrote the script (attachment) to make csv files (attachment).  </p>",
      "votes": null,
      "replies": [
        {
          "id": 171687,
          "author_name": "wendykan",
          "author_url": "",
          "post_date": "03/31/2017 02:42:51",
          "content": "<p>Thanks for spotting this! This is fixed now. Files are now postfixed with <code>_v2</code>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 171284,
      "author_name": "yadavsarthak",
      "author_url": "",
      "post_date": "03/29/2017 09:40:27",
      "content": "<p>has the data been updated on kaggle kernels?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 171664,
      "author_name": "hanyangkyushu",
      "author_url": "",
      "post_date": "03/30/2017 23:30:40",
      "content": "<p>I think there is something wrong in these new additional data and fixed label file. </p>\n\n<p>For example, the fixed_label.csv tells me 10.jpg has to be type_1, but it is still in additional_Type_2.</p>\n\n<p>Also, as written by Kambarakun, there are some duplication in file names.</p>\n\n<p>It should be fixed, I think.</p>",
      "votes": null,
      "replies": [
        {
          "id": 171688,
          "author_name": "wendykan",
          "author_url": "",
          "post_date": "03/31/2017 02:43:02",
          "content": "<p>Thanks for spotting this! This is fixed now. Files are now postfixed with <code>_v2</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 171695,
          "author_name": "hanyangkyushu",
          "author_url": "",
          "post_date": "03/31/2017 03:18:02",
          "content": "<p>Thank you so much for quick correction~!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 172306,
      "author_name": "akumar6",
      "author_url": "",
      "post_date": "04/03/2017 07:06:05",
      "content": "<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 172376,
      "author_name": "zorancupic",
      "author_url": "",
      "post_date": "04/03/2017 13:57:14",
      "content": "<p>I think there is still a discrepancy in the data. I've downloaded the <strong>train</strong> data/images and the fixed_labels_v2 on 2nd April 2017. For example 1001.jpg is still available in the 'Type_2' folder but the new label should be Type_1 , or do some images have the same id in train and additional data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 172472,
          "author_name": "wendykan",
          "author_url": "",
          "post_date": "04/03/2017 20:51:32",
          "content": "<p>Correct - images have the same ids in train/test and additional since they all start from 0.jpg. All the fixed_labels are on the additional dataset, not train. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 173836,
      "author_name": "jessiecarnegie7777",
      "author_url": "",
      "post_date": "04/08/2017 22:38:46",
      "content": "<p>Can we please get a torrent somehow? I've downloaded the data twice and I'm still getting corrupted files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 174248,
      "author_name": "pachinko",
      "author_url": "",
      "post_date": "04/10/2017 19:56:09",
      "content": "<p>Hello anybody wanna try out the torrent file I included?\nI have never tried to create a torrent before.</p>\n\n<p>If this is working, I'll try to create more(or a single one) torrent after I succeed in downloading all the big zip files tomorrow.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 175144,
      "author_name": "pachinko",
      "author_url": "",
      "post_date": "04/14/2017 01:19:38",
      "content": "<p>Hi everybody, I've make these 5 big image zip files into a torrent.\nI am not sure if hosting this will block my internet or something.\nAnyone who hasn't successfully downloaded could give it a try.</p>",
      "votes": null,
      "replies": [
        {
          "id": 188554,
          "author_name": "icanfly",
          "author_url": "",
          "post_date": "06/03/2017 07:24:03",
          "content": "<p>Hi, Does the torrent file still work? It seems to fail.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "170904": "Hi all,\n\nAs described in [this][1] thread, we had fixed the labeling errors in the additional dataset. \n\n1. If you had downloaded additional.7z file before March 24, 2017, you can use two csv files to avoid re-downloading. fixed_labels has the file names, their old labels and new labels. removed_files has the list of files to remove. \n\n2. If you haven't downloaded the additional dataset, they are now available (with the correct labels and no duplicates) as 3 parts: additional_Type_{x}.7z, x=1,2,3. \n\n3. We don't officially support the torrent download for this competition, so I have pulled down the torrent file. If someone in the community wants to host, please send me the torrent file and I'll upload.\n\n4. Both the Intel server and Kaggle Kernels will have the updated additional data soon. \n\nThank you all for your patience. \n\nKaggle admin\n\n\n  [1]: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30621",
    "170908": "How to wget those datasets on an ec2 instance? I tried:\n\n`wget https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/download/train.7z`\n\nbut it doesn't download actual dataset",
    "170933": "You may try this:\n\n```\nwget https://kaggle2.blob.core.windows.net/competitions-data/kaggle/6243/train.7z?sv=2015-12-11&sr=b&sig=LuJRoYHig6Df88fnn5SaXFVO6RjT8O6TswVfD%2FrE46M%3D&se=2017-03-31T02%3A12%3A39Z&sp=r\n```",
    "170946": "It gives me 404 error",
    "170981": "thank you",
    "170983": "Thanks so much! 😀  \nI want to announce once more, when you have done updating datasets (#4.).",
    "171099": "you can use kaggle-cli\nhttps://github.com/floydwch/kaggle-cli",
    "171133": "I tried to extract 7z files on Colfax Cluster, but additional_Type_2.7z are broken.\n\nI'm downloading at [web page of kaggle](https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/data).  \nI hope web page's file is not broken:)\n\n----------\n$ md5sum /data/kaggle_3.27/additional_Type_2.7z\n\n38bf446a2b66a2d20fc7a4c6d9089131 /data/kaggle_3.27/additional_Type_2.7z\n\n$ 7za x /data/kaggle_3.27/additional_Type_2.7z\n\n7-Zip (a) [64] 15.09 beta : Copyright (c) 1999-2015 Igor Pavlov : 2015-10-16\np7zip Version 15.09 beta (locale=ja_JP.UTF-8,Utf16=on,HugeFiles=on,64 bits,8 CPUs Intel Core Processor (Broadwell) (306D2),ASM,AES-NI)\n\nScanning the drive for archives:\n1 file, 18491891712 bytes (18 GiB)\n\nExtracting archive: /data/kaggle_3.27/additional_Type_2.7z\nERROR: /data/kaggle_3.27/additional_Type_2.7z\n/data/kaggle_3.27/additional_Type_2.7z\nOpen ERROR: Can not open the file as [7z] archive\n\n\nERRORS:\nHeaders Error\nWARNINGS:\nThere are data after the end of archive\n    \nCan't open as archive: 1\nFiles: 0\nSize:       0\nCompressed: 0",
    "171141": "Cool! thank you that works!",
    "171175": "I have extracted additional_Type_2.7z without error, downloaded on kaggle's web page.  \nSorry for irregular way.\n\n* on kaggle's web page  \n  filesize = 15871727173  \n  md5sum = ee6c378c086bb3e77cda29b1bf4828d6  \n  sha1sum = 4495bf9e7de477ba64daf01e3be6e04ecbc6151a\n\n* on Colfax Cluster (broken)  \n  filesize = 18491891712  \n  md5sum = 38bf446a2b66a2d20fc7a4c6d9089131  \n  sha1sum = 1721366bea847c2dcd1a84cbe84453b51695daad",
    "171184": "Thanks for point out the additional_Type_2.7z with broken checksum on Colfax. We have fixed it now.",
    "171254": "Sorry to bother you again.  \n**I found duplication on removed_files.csv and fixed_labels.csv files.**  \n\nDuplicate Remove  \n./Type_3/6821.jpg  \n\nRemove or Fix?  \n./Type_2/1018.jpg, ./Type_2/1094.jpg, ./Type_2/114.jpg,  ./Type_3/1295.jpg, ./Type_2/1319.jpg,  \n./Type_2/1510.jpg, ./Type_3/1762.jpg, ./Type_2/1799.jpg, ./Type_3/1822.jpg, ./Type_1/1825.jpg,  \n./Type_2/2010.jpg, ./Type_1/2151.jpg, ./Type_1/2320.jpg, ./Type_1/2414.jpg, ./Type_3/2491.jpg,  \n./Type_3/2511.jpg, ./Type_2/2562.jpg, ./Type_3/2613.jpg, ./Type_2/2663.jpg, ./Type_2/2860.jpg,  \n./Type_2/2937.jpg, ./Type_3/2998.jpg, ./Type_2/3017.jpg, ./Type_3/304.jpg,  ./Type_1/3355.jpg,  \n./Type_2/3615.jpg, ./Type_1/3662.jpg, ./Type_3/3692.jpg, ./Type_3/3849.jpg, ./Type_3/3856.jpg,  \n./Type_1/3897.jpg, ./Type_3/423.jpg,  ./Type_2/474.jpg,  ./Type_2/502.jpg,  ./Type_2/664.jpg,  \n./Type_2/68.jpg,   ./Type_2/811.jpg,  ./Type_3/821.jpg,  ./Type_2/989.jpg  \n\nSo, I check the number of additional images  \n(removed_and_fixed means removed -> fixed, removed_and_fixed means fixed -> removed).  \nI wrote the rough script (attachment) on python 2.7.  \n**This script is destructive.  \nPlease backup original additional directory!**  \n\nname,Type_1,Type_2,Type_3,total  \nadditional_old,   1192,3619,2113,6924  \nkaggle_kernel_ago,1192,3619,2113,6924  \ncolfax_old,       1192,3619,2113,6924  \nre-download,      1125,3617,2030,6772  \nkaggle_kernel_now,1125,3617,2030,6772  \ncolfax_now,       NaN, NaN, NaN, NaN  \nremoved_and_fixed,1189,3565,1974,6728  \nfixed_and_removed,1201,3580,1986,6767  \n\nI can't check /data/kaggle_3.27/Type_* on Colfax Cluster, because of permission denied.  \n\n/data/kaggle_3.27  \ndrwxr-x---. 2 root root 32768  3月 27 14:35 Type_1  \ndrwxr-x---. 2 root root 73728  3月 27 14:35 Type_2  \ndrwxr-x---. 2 root root 49152  3月 27 14:35 Type_3  \n\n/data/kaggle/additional  \ndrwxr-xr-x. 2 root root 36864  3月  8 15:15 Type_1  \ndrwxr-xr-x. 2 root root 77824  3月  8 15:15 Type_2  \ndrwxr-xr-x. 2 root root 53248  3月  8 15:15 Type_3  \n\n**I think removed_files.csv and fixed_labels.csv are wrong, because of duplication.**  \nI think it is good to use re-download dataset.  \n**Any time suits, please announce us which is the correct dataset:)**  \n\nADD: I do not mean to be rude, but I wrote the script (attachment) to make csv files (attachment).",
    "171284": "has the data been updated on kaggle kernels?",
    "171664": "I think there is something wrong in these new additional data and fixed label file. \n\nFor example, the fixed_label.csv tells me 10.jpg has to be type_1, but it is still in additional_Type_2.\n\nAlso, as written by Kambarakun, there are some duplication in file names.\n\nIt should be fixed, I think.",
    "171687": "Thanks for spotting this! This is fixed now. Files are now postfixed with `_v2`.",
    "171688": "Thanks for spotting this! This is fixed now. Files are now postfixed with `_v2`.",
    "171695": "Thank you so much for quick correction~!",
    "172306": "Thanks.",
    "172376": "I think there is still a discrepancy in the data. I've downloaded the **train** data/images and the fixed_labels_v2 on 2nd April 2017. For example 1001.jpg is still available in the 'Type_2' folder but the new label should be Type_1 , or do some images have the same id in train and additional data?",
    "172472": "Correct - images have the same ids in train/test and additional since they all start from 0.jpg. All the fixed_labels are on the additional dataset, not train.",
    "173836": "Can we please get a torrent somehow? I've downloaded the data twice and I'm still getting corrupted files.",
    "174248": "Hello anybody wanna try out the torrent file I included?\nI have never tried to create a torrent before.\n\nIf this is working, I'll try to create more(or a single one) torrent after I succeed in downloading all the big zip files tomorrow.",
    "175144": "Hi everybody, I've make these 5 big image zip files into a torrent.\nI am not sure if hosting this will block my internet or something.\nAnyone who hasn't successfully downloaded could give it a try.",
    "188554": "Hi, Does the torrent file still work? It seems to fail."
  },
  "source": "meta"
}