{
  "id": 40086,
  "title": "Hello from Cdiscount",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40086",
  "author_name": "",
  "post_date": "2017-09-27T16:19:12.218005200Z",
  "votes": 24,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Hi everyone. First, we would like to thank you all for participating to the competition. We are already impressed by your performances: we did not believe the 70% mark would be hit in less than two weeks.</p>\n\n<p>As said in the description, the algorithms will have an immediate impact on our customers experience. For example, no one likes getting a page polluted by low-price, badly classified accessories after sorting the products by price. Beyond that, we hope that the data set we release will prove of valuable help to researchers, educators and students to improve image classification algorithms in an extreme multi-class context. Well, it’s not every day 15 million labeled images are released.</p>\n\n<p>Feel free to post in this thread your questions, comments, critics and any kind of feedback!</p>\n\n<p>Cdiscount data science team</p>\n\n<p>PS: you can contact us directly at datascience@cdiscount.com</p>",
  "messages": [
    {
      "id": "224770",
      "postDate": "09/27/2017 16:19:12",
      "content": "<p>Hi everyone. First, we would like to thank you all for participating to the competition. We are already impressed by your performances: we did not believe the 70% mark would be hit in less than two weeks.</p>\n\n<p>As said in the description, the algorithms will have an immediate impact on our customers experience. For example, no one likes getting a page polluted by low-price, badly classified accessories after sorting the products by price. Beyond that, we hope that the data set we release will prove of valuable help to researchers, educators and students to improve image classification algorithms in an extreme multi-class context. Well, it’s not every day 15 million labeled images are released.</p>\n\n<p>Feel free to post in this thread your questions, comments, critics and any kind of feedback!</p>\n\n<p>Cdiscount data science team</p>\n\n<p>PS: you can contact us directly at datascience@cdiscount.com</p>",
      "rawMarkdown": "Hi everyone. First, we would like to thank you all for participating to the competition. We are already impressed by your performances: we did not believe the 70% mark would be hit in less than two weeks.\n \nAs said in the description, the algorithms will have an immediate impact on our customers experience. For example, no one likes getting a page polluted by low-price, badly classified accessories after sorting the products by price. Beyond that, we hope that the data set we release will prove of valuable help to researchers, educators and students to improve image classification algorithms in an extreme multi-class context. Well, it’s not every day 15 million labeled images are released.\n \nFeel free to post in this thread your questions, comments, critics and any kind of feedback!\n \nCdiscount data science team\n\nPS: you can contact us directly at datascience@cdiscount.com",
      "votes": null
    },
    {
      "id": "224784",
      "postDate": "09/27/2017 16:46:33",
      "content": "<blockquote>\n  <p>we did not believe the 70% mark would be hit in less than two weeks</p>\n</blockquote>\n\n<p>What is your score?</p>",
      "rawMarkdown": "&gt; we did not believe the 70% mark would be hit in less than two weeks\n\nWhat is your score?",
      "votes": null
    },
    {
      "id": "224801",
      "postDate": "09/27/2017 17:32:25",
      "content": "<p>Can you provide us AWS credits for people who wants to enter this competition, but that does not have sufficient hardware to process 60+GB of data?</p>",
      "rawMarkdown": "Can you provide us AWS credits for people who wants to enter this competition, but that does not have sufficient hardware to process 60+GB of data?",
      "votes": null
    },
    {
      "id": "224945",
      "postDate": "09/28/2017 00:27:52",
      "content": "<p>Can you provide a torrent or some other way to download the data. Downloading the training data is taking too much time </p>",
      "rawMarkdown": "Can you provide a torrent or some other way to download the data. Downloading the training data is taking too much time",
      "votes": null
    },
    {
      "id": "225015",
      "postDate": "09/28/2017 03:51:07",
      "content": "<p>How do you think deep learning(vision only) is going to disrupt e-commerce space ?.  Also are you planning to use any natural language processing to display appropriate images from their descriptions ?</p>",
      "rawMarkdown": "How do you think deep learning(vision only) is going to disrupt e-commerce space ?.  Also are you planning to use any natural language processing to display appropriate images from their descriptions ?",
      "votes": null
    },
    {
      "id": "225097",
      "postDate": "09/28/2017 08:40:29",
      "content": "<p>Not planned for this challenge, unfortunately. But definitively something we will negotiate next time we organize a competition. Though probably not on Amazon, a direct competitor!</p>",
      "rawMarkdown": "Not planned for this challenge, unfortunately. But definitively something we will negotiate next time we organize a competition. Though probably not on Amazon, a direct competitor!",
      "votes": null
    },
    {
      "id": "225098",
      "postDate": "09/28/2017 08:41:06",
      "content": "<p>We're around 70%, on a different training/testing set however</p>",
      "rawMarkdown": "We're around 70%, on a different training/testing set however",
      "votes": null
    },
    {
      "id": "225099",
      "postDate": "09/28/2017 08:43:38",
      "content": "<p>We're looking at this with Kaggle</p>",
      "rawMarkdown": "We're looking at this with Kaggle",
      "votes": null
    },
    {
      "id": "225127",
      "postDate": "09/28/2017 10:19:14",
      "content": "<p>How do you handle products which do not belong to any of the listed categories?\nHow do you detect all those products?</p>",
      "rawMarkdown": "How do you handle products which do not belong to any of the listed categories?\nHow do you detect all those products?",
      "votes": null
    },
    {
      "id": "225371",
      "postDate": "09/28/2017 21:19:03",
      "content": "<p>Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!</p>",
      "rawMarkdown": "Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!",
      "votes": null
    },
    {
      "id": "225385",
      "postDate": "09/28/2017 21:40:24",
      "content": "<p>What is your score for text classify?</p>",
      "rawMarkdown": "What is your score for text classify?",
      "votes": null
    },
    {
      "id": "225390",
      "postDate": "09/28/2017 22:50:11",
      "content": "<p>Do you have any runtime requirements for productionizing a solution? Given that many kaggle competitions are won by unholy ensembles with compute costs upwards of 10^10 FLOPs, is this a barrier to deployment or is your use-case sufficiently sparse that it's worth it for the accuracy gains?</p>",
      "rawMarkdown": "Do you have any runtime requirements for productionizing a solution? Given that many kaggle competitions are won by unholy ensembles with compute costs upwards of 10^10 FLOPs, is this a barrier to deployment or is your use-case sufficiently sparse that it's worth it for the accuracy gains?",
      "votes": null
    },
    {
      "id": "225490",
      "postDate": "09/29/2017 08:32:59",
      "content": "<p>great, thanks a lot!</p>",
      "rawMarkdown": "great, thanks a lot!",
      "votes": null
    },
    {
      "id": "225491",
      "postDate": "09/29/2017 08:40:27",
      "content": "<p>Tough question... :-) Deep learning may have a great potential in a context of specialized vendors. Style recommendation for a fashion vendor, for example. But such companies may not have the financial or technical capabilities to invest in deep learning. In our case it serves to classify products: great for users experience, but disruptive in terms of e-commerce space would be quite strong a statement.</p>\n\n<p>We do not need to associate relevant images to a given description. We do use text-based classification algorithms, though.</p>",
      "rawMarkdown": "Tough question... :-) Deep learning may have a great potential in a context of specialized vendors. Style recommendation for a fashion vendor, for example. But such companies may not have the financial or technical capabilities to invest in deep learning. In our case it serves to classify products: great for users experience, but disruptive in terms of e-commerce space would be quite strong a statement.\n\nWe do not need to associate relevant images to a given description. We do use text-based classification algorithms, though.",
      "votes": null
    },
    {
      "id": "225494",
      "postDate": "09/29/2017 08:45:36",
      "content": "<p>The classification algorithm gives predictions with low associated probabilities on such products, so the predictions are rejected and the classification is done manually. This may involve the creation of new categories.</p>",
      "rawMarkdown": "The classification algorithm gives predictions with low associated probabilities on such products, so the predictions are rejected and the classification is done manually. This may involve the creation of new categories.",
      "votes": null
    },
    {
      "id": "225503",
      "postDate": "09/29/2017 09:08:21",
      "content": "<p>In production, we set up a threshold value for the level of confidence (i.e. the probability) of the predictions, in order to guarantee an error rate lower that some imposed value. We are able to classify two thirds of the products with the imposed level of accuracy.</p>",
      "rawMarkdown": "In production, we set up a threshold value for the level of confidence (i.e. the probability) of the predictions, in order to guarantee an error rate lower that some imposed value. We are able to classify two thirds of the products with the imposed level of accuracy.",
      "votes": null
    },
    {
      "id": "225506",
      "postDate": "09/29/2017 09:10:56",
      "content": "<p>Yes there is a trade-off between method accuracy and production constraints. It is unlikely that the winning algorithms will be deployed as is. We will rather try to understand the general philosophy and bags of \"tricks\" used by the winners, in order to guide the development of our own solution.</p>",
      "rawMarkdown": "Yes there is a trade-off between method accuracy and production constraints. It is unlikely that the winning algorithms will be deployed as is. We will rather try to understand the general philosophy and bags of \"tricks\" used by the winners, in order to guide the development of our own solution.",
      "votes": null
    },
    {
      "id": "226642",
      "postDate": "10/02/2017 21:42:53",
      "content": "<p>Train and Test dataset is just dump from database or something special crafted for Kaggle?</p>",
      "rawMarkdown": "Train and Test dataset is just dump from database or something special crafted for Kaggle?",
      "votes": null
    },
    {
      "id": "226876",
      "postDate": "10/03/2017 08:35:28",
      "content": "<p>Hi, yes these are real-world data sets, from our current catalogue of products</p>",
      "rawMarkdown": "Hi, yes these are real-world data sets, from our current catalogue of products",
      "votes": null
    },
    {
      "id": "226992",
      "postDate": "10/03/2017 13:53:11",
      "content": "<p>How I can contact Cdiscount directly? I have few related analyses that I would like to share with Cdiscount Tech team private.</p>",
      "rawMarkdown": "How I can contact Cdiscount directly? I have few related analyses that I would like to share with Cdiscount Tech team private.",
      "votes": null
    },
    {
      "id": "227369",
      "postDate": "10/04/2017 07:54:12",
      "content": "<p>we added a contact e-mail in the welcome post above</p>",
      "rawMarkdown": "we added a contact e-mail in the welcome post above",
      "votes": null
    },
    {
      "id": "227566",
      "postDate": "10/04/2017 17:19:46",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    },
    {
      "id": "227950",
      "postDate": "10/05/2017 14:24:38",
      "content": "<p>If you can answer the following questions that will help me in the research paper:</p>\n\n<ol>\n<li><p>Can I know what machine learning algorithm did you use for text description classifier one?</p></li>\n<li><p>What was the accuracy percentage?</p></li>\n<li><p>Have you tried any image classifier before posting the project in Kaggle?</p></li>\n</ol>\n\n<p>If you can give details</p>",
      "rawMarkdown": "If you can answer the following questions that will help me in the research paper:\n\n1. Can I know what machine learning algorithm did you use for text description classifier one?\n\n2. What was the accuracy percentage?\n\n3. Have you tried any image classifier before posting the project in Kaggle?\n\n\nIf you can give details",
      "votes": null
    },
    {
      "id": "228252",
      "postDate": "10/06/2017 08:00:38",
      "content": "<p>Hi Wedad here are the answers to your questions:</p>\n\n<ol>\n<li>We currently use KNN over tf-idf on the title &amp; description of the items (with a lemmatisation to prepare the data)</li>\n<li>In production, we classify products only on categories where we can reach more than 90% of accuracy. Given this constraint, we are able to classify more than 2/3 of the catalog today.</li>\n<li>Yes, we have been working on classifying products based on their images - best results so far obtained using Inception v3.</li>\n</ol>",
      "rawMarkdown": "Hi Wedad here are the answers to your questions:\n\n1. We currently use KNN over tf-idf on the title &amp; description of the items (with a lemmatisation to prepare the data)\n2. In production, we classify products only on categories where we can reach more than 90% of accuracy. Given this constraint, we are able to classify more than 2/3 of the catalog today.\n3. Yes, we have been working on classifying products based on their images - best results so far obtained using Inception v3.",
      "votes": null
    },
    {
      "id": "242688",
      "postDate": "11/12/2017 11:40:08",
      "content": "<p>I do not have 60GB+ of space to use in my pc. I do have GPUs and RAM to use though, any ideas on how to get access to the data?</p>",
      "rawMarkdown": "I do not have 60GB+ of space to use in my pc. I do have GPUs and RAM to use though, any ideas on how to get access to the data?",
      "votes": null
    },
    {
      "id": "242695",
      "postDate": "11/12/2017 12:02:30",
      "content": "<p>I recommend that you buy additional strage.<br>\n500GB would be enough (SSD is better).</p>",
      "rawMarkdown": "I recommend that you buy additional strage.<br>\n500GB would be enough (SSD is better).",
      "votes": null
    },
    {
      "id": "242696",
      "postDate": "11/12/2017 12:03:58",
      "content": "<p>What is the size of this dataset? Maybe I will store it somewhere else and access by IP</p>",
      "rawMarkdown": "What is the size of this dataset? Maybe I will store it somewhere else and access by IP",
      "votes": null
    },
    {
      "id": "327938",
      "postDate": "05/13/2018 01:07:36",
      "content": "<p>Hi, CdiscountDS. Could you please share the data(.bson or .jpg)? The data  now has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thanks~</p>",
      "rawMarkdown": "Hi, CdiscountDS. Could you please share the data(.bson or .jpg)? The data  now has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thanks~",
      "votes": null
    },
    {
      "id": "432684",
      "postDate": "12/04/2018 07:24:18",
      "content": "<p>I want to get the dataset...</p>",
      "rawMarkdown": "I want to get the dataset...",
      "votes": null
    },
    {
      "id": "432742",
      "postDate": "12/04/2018 09:14:38",
      "content": "<p>Hi! Have you got the dataset? I am a beginner in Product classification and I also want to use the dataset <br>\n to reproduce some experimental results.</p>",
      "rawMarkdown": "Hi! Have you got the dataset? I am a beginner in Product classification and I also want to use the dataset  \n to reproduce some experimental results.",
      "votes": null
    },
    {
      "id": "495407",
      "postDate": "03/21/2019 03:42:52",
      "content": "<p>Hi inversion, are these torrent files still available, even after the competition closed? I'm asking since I couldn't find them in the Data tab. Thanks in advance!</p>",
      "rawMarkdown": "Hi inversion, are these torrent files still available, even after the competition closed? I'm asking since I couldn't find them in the Data tab. Thanks in advance!",
      "votes": null
    },
    {
      "id": "1168499",
      "postDate": "01/25/2021 02:02:09",
      "content": "<p>could this dataset be used for commercial purpose?</p>",
      "rawMarkdown": "could this dataset be used for commercial purpose?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1168499,
      "author_name": "zhihaneb",
      "author_url": "",
      "post_date": "01/25/2021 02:02:09",
      "content": "<p>could this dataset be used for commercial purpose?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224784,
      "author_name": "radustoicescu",
      "author_url": "",
      "post_date": "09/27/2017 16:46:33",
      "content": "<blockquote>\n  <p>we did not believe the 70% mark would be hit in less than two weeks</p>\n</blockquote>\n\n<p>What is your score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225098,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/28/2017 08:41:06",
          "content": "<p>We're around 70%, on a different training/testing set however</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 224801,
      "author_name": "thomasseleck",
      "author_url": "",
      "post_date": "09/27/2017 17:32:25",
      "content": "<p>Can you provide us AWS credits for people who wants to enter this competition, but that does not have sufficient hardware to process 60+GB of data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225097,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/28/2017 08:40:29",
          "content": "<p>Not planned for this challenge, unfortunately. But definitively something we will negotiate next time we organize a competition. Though probably not on Amazon, a direct competitor!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 224945,
      "author_name": "abhimathur84",
      "author_url": "",
      "post_date": "09/28/2017 00:27:52",
      "content": "<p>Can you provide a torrent or some other way to download the data. Downloading the training data is taking too much time </p>",
      "votes": null,
      "replies": [
        {
          "id": 225099,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/28/2017 08:43:38",
          "content": "<p>We're looking at this with Kaggle</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225371,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "09/28/2017 21:19:03",
          "content": "<p>Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225490,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/29/2017 08:32:59",
          "content": "<p>great, thanks a lot!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 495407,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "03/21/2019 03:42:52",
          "content": "<p>Hi inversion, are these torrent files still available, even after the competition closed? I'm asking since I couldn't find them in the Data tab. Thanks in advance!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225015,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "09/28/2017 03:51:07",
      "content": "<p>How do you think deep learning(vision only) is going to disrupt e-commerce space ?.  Also are you planning to use any natural language processing to display appropriate images from their descriptions ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225491,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/29/2017 08:40:27",
          "content": "<p>Tough question... :-) Deep learning may have a great potential in a context of specialized vendors. Style recommendation for a fashion vendor, for example. But such companies may not have the financial or technical capabilities to invest in deep learning. In our case it serves to classify products: great for users experience, but disruptive in terms of e-commerce space would be quite strong a statement.</p>\n\n<p>We do not need to associate relevant images to a given description. We do use text-based classification algorithms, though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225127,
      "author_name": "andreaslup",
      "author_url": "",
      "post_date": "09/28/2017 10:19:14",
      "content": "<p>How do you handle products which do not belong to any of the listed categories?\nHow do you detect all those products?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225494,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/29/2017 08:45:36",
          "content": "<p>The classification algorithm gives predictions with low associated probabilities on such products, so the predictions are rejected and the classification is done manually. This may involve the creation of new categories.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225385,
      "author_name": "patron3301",
      "author_url": "",
      "post_date": "09/28/2017 21:40:24",
      "content": "<p>What is your score for text classify?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225503,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/29/2017 09:08:21",
          "content": "<p>In production, we set up a threshold value for the level of confidence (i.e. the probability) of the predictions, in order to guarantee an error rate lower that some imposed value. We are able to classify two thirds of the products with the imposed level of accuracy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225390,
      "author_name": "ajmooch",
      "author_url": "",
      "post_date": "09/28/2017 22:50:11",
      "content": "<p>Do you have any runtime requirements for productionizing a solution? Given that many kaggle competitions are won by unholy ensembles with compute costs upwards of 10^10 FLOPs, is this a barrier to deployment or is your use-case sufficiently sparse that it's worth it for the accuracy gains?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225506,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "09/29/2017 09:10:56",
          "content": "<p>Yes there is a trade-off between method accuracy and production constraints. It is unlikely that the winning algorithms will be deployed as is. We will rather try to understand the general philosophy and bags of \"tricks\" used by the winners, in order to guide the development of our own solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 226642,
      "author_name": "patron3301",
      "author_url": "",
      "post_date": "10/02/2017 21:42:53",
      "content": "<p>Train and Test dataset is just dump from database or something special crafted for Kaggle?</p>",
      "votes": null,
      "replies": [
        {
          "id": 226876,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "10/03/2017 08:35:28",
          "content": "<p>Hi, yes these are real-world data sets, from our current catalogue of products</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 226992,
      "author_name": "patron3301",
      "author_url": "",
      "post_date": "10/03/2017 13:53:11",
      "content": "<p>How I can contact Cdiscount directly? I have few related analyses that I would like to share with Cdiscount Tech team private.</p>",
      "votes": null,
      "replies": [
        {
          "id": 227369,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "10/04/2017 07:54:12",
          "content": "<p>we added a contact e-mail in the welcome post above</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 227566,
          "author_name": "patron3301",
          "author_url": "",
          "post_date": "10/04/2017 17:19:46",
          "content": "<p>Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 227950,
      "author_name": "wedadanbtawi",
      "author_url": "",
      "post_date": "10/05/2017 14:24:38",
      "content": "<p>If you can answer the following questions that will help me in the research paper:</p>\n\n<ol>\n<li><p>Can I know what machine learning algorithm did you use for text description classifier one?</p></li>\n<li><p>What was the accuracy percentage?</p></li>\n<li><p>Have you tried any image classifier before posting the project in Kaggle?</p></li>\n</ol>\n\n<p>If you can give details</p>",
      "votes": null,
      "replies": [
        {
          "id": 228252,
          "author_name": "cdiscountds",
          "author_url": "",
          "post_date": "10/06/2017 08:00:38",
          "content": "<p>Hi Wedad here are the answers to your questions:</p>\n\n<ol>\n<li>We currently use KNN over tf-idf on the title &amp; description of the items (with a lemmatisation to prepare the data)</li>\n<li>In production, we classify products only on categories where we can reach more than 90% of accuracy. Given this constraint, we are able to classify more than 2/3 of the catalog today.</li>\n<li>Yes, we have been working on classifying products based on their images - best results so far obtained using Inception v3.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 242688,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "11/12/2017 11:40:08",
      "content": "<p>I do not have 60GB+ of space to use in my pc. I do have GPUs and RAM to use though, any ideas on how to get access to the data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 242695,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "11/12/2017 12:02:30",
          "content": "<p>I recommend that you buy additional strage.<br>\n500GB would be enough (SSD is better).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242696,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "11/12/2017 12:03:58",
          "content": "<p>What is the size of this dataset? Maybe I will store it somewhere else and access by IP</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327938,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "05/13/2018 01:07:36",
      "content": "<p>Hi, CdiscountDS. Could you please share the data(.bson or .jpg)? The data  now has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thanks~</p>",
      "votes": null,
      "replies": [
        {
          "id": 432742,
          "author_name": "alexuan",
          "author_url": "",
          "post_date": "12/04/2018 09:14:38",
          "content": "<p>Hi! Have you got the dataset? I am a beginner in Product classification and I also want to use the dataset <br>\n to reproduce some experimental results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 432684,
      "author_name": "sjtuqin",
      "author_url": "",
      "post_date": "12/04/2018 07:24:18",
      "content": "<p>I want to get the dataset...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "224770": "Hi everyone. First, we would like to thank you all for participating to the competition. We are already impressed by your performances: we did not believe the 70% mark would be hit in less than two weeks.\n \nAs said in the description, the algorithms will have an immediate impact on our customers experience. For example, no one likes getting a page polluted by low-price, badly classified accessories after sorting the products by price. Beyond that, we hope that the data set we release will prove of valuable help to researchers, educators and students to improve image classification algorithms in an extreme multi-class context. Well, it’s not every day 15 million labeled images are released.\n \nFeel free to post in this thread your questions, comments, critics and any kind of feedback!\n \nCdiscount data science team\n\nPS: you can contact us directly at datascience@cdiscount.com",
    "224784": "&gt; we did not believe the 70% mark would be hit in less than two weeks\n\nWhat is your score?",
    "224801": "Can you provide us AWS credits for people who wants to enter this competition, but that does not have sufficient hardware to process 60+GB of data?",
    "224945": "Can you provide a torrent or some other way to download the data. Downloading the training data is taking too much time",
    "225015": "How do you think deep learning(vision only) is going to disrupt e-commerce space ?.  Also are you planning to use any natural language processing to display appropriate images from their descriptions ?",
    "225097": "Not planned for this challenge, unfortunately. But definitively something we will negotiate next time we organize a competition. Though probably not on Amazon, a direct competitor!",
    "225098": "We're around 70%, on a different training/testing set however",
    "225099": "We're looking at this with Kaggle",
    "225127": "How do you handle products which do not belong to any of the listed categories?\nHow do you detect all those products?",
    "225371": "Torrent files for train and test are now available on the data page. Please let me know how they work out for you. Thanks!",
    "225385": "What is your score for text classify?",
    "225390": "Do you have any runtime requirements for productionizing a solution? Given that many kaggle competitions are won by unholy ensembles with compute costs upwards of 10^10 FLOPs, is this a barrier to deployment or is your use-case sufficiently sparse that it's worth it for the accuracy gains?",
    "225490": "great, thanks a lot!",
    "225491": "Tough question... :-) Deep learning may have a great potential in a context of specialized vendors. Style recommendation for a fashion vendor, for example. But such companies may not have the financial or technical capabilities to invest in deep learning. In our case it serves to classify products: great for users experience, but disruptive in terms of e-commerce space would be quite strong a statement.\n\nWe do not need to associate relevant images to a given description. We do use text-based classification algorithms, though.",
    "225494": "The classification algorithm gives predictions with low associated probabilities on such products, so the predictions are rejected and the classification is done manually. This may involve the creation of new categories.",
    "225503": "In production, we set up a threshold value for the level of confidence (i.e. the probability) of the predictions, in order to guarantee an error rate lower that some imposed value. We are able to classify two thirds of the products with the imposed level of accuracy.",
    "225506": "Yes there is a trade-off between method accuracy and production constraints. It is unlikely that the winning algorithms will be deployed as is. We will rather try to understand the general philosophy and bags of \"tricks\" used by the winners, in order to guide the development of our own solution.",
    "226642": "Train and Test dataset is just dump from database or something special crafted for Kaggle?",
    "226876": "Hi, yes these are real-world data sets, from our current catalogue of products",
    "226992": "How I can contact Cdiscount directly? I have few related analyses that I would like to share with Cdiscount Tech team private.",
    "227369": "we added a contact e-mail in the welcome post above",
    "227566": "Thank you.",
    "227950": "If you can answer the following questions that will help me in the research paper:\n\n1. Can I know what machine learning algorithm did you use for text description classifier one?\n\n2. What was the accuracy percentage?\n\n3. Have you tried any image classifier before posting the project in Kaggle?\n\n\nIf you can give details",
    "228252": "Hi Wedad here are the answers to your questions:\n\n1. We currently use KNN over tf-idf on the title &amp; description of the items (with a lemmatisation to prepare the data)\n2. In production, we classify products only on categories where we can reach more than 90% of accuracy. Given this constraint, we are able to classify more than 2/3 of the catalog today.\n3. Yes, we have been working on classifying products based on their images - best results so far obtained using Inception v3.",
    "242688": "I do not have 60GB+ of space to use in my pc. I do have GPUs and RAM to use though, any ideas on how to get access to the data?",
    "242695": "I recommend that you buy additional strage.<br>\n500GB would be enough (SSD is better).",
    "242696": "What is the size of this dataset? Maybe I will store it somewhere else and access by IP",
    "327938": "Hi, CdiscountDS. Could you please share the data(.bson or .jpg)? The data  now has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thanks~",
    "432684": "I want to get the dataset...",
    "432742": "Hi! Have you got the dataset? I am a beginner in Product classification and I also want to use the dataset  \n to reproduce some experimental results.",
    "495407": "Hi inversion, are these torrent files still available, even after the competition closed? I'm asking since I couldn't find them in the Data tab. Thanks in advance!",
    "1168499": "could this dataset be used for commercial purpose?"
  },
  "source": "meta"
}