{
  "id": 240187,
  "title": "External data - 1000 triple-read Covid-19 X-rays",
  "url": "/competitions/siim-covid19-detection/discussion/240187",
  "author_name": "raddar",
  "post_date": "2021-05-18T19:37:03.052000",
  "votes": 111,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I have shared a very useful dataset, which contains 1000 X-rays, which were rated by 3 different radiologists using the same taxonomy (<code>\"negative\", \"typical\", \"indeterminate\", \"atypical\"</code>) as in the competition.</p>\n<p><a href=\"https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests\" target=\"_blank\">https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests</a><br>\nPlease use it wisely!</p>\n<p>If you looked at this dataset more closely, only 600 out of 1000 cases have radiologists unanimously agree about Covid presence in X-ray- this is insanely low agreement!</p>\n<p>On top of that, there are 150 cases where all radiologists think that there is no Covid (although tests have it confirmed)</p>",
  "messages": [
    {
      "id": 1313908,
      "postDate": "2021-05-18T19:37:03.053Z",
      "content": "<p>I have shared a very useful dataset, which contains 1000 X-rays, which were rated by 3 different radiologists using the same taxonomy (<code>\"negative\", \"typical\", \"indeterminate\", \"atypical\"</code>) as in the competition.</p>\n<p><a href=\"https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests\" target=\"_blank\">https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests</a><br>\nPlease use it wisely!</p>\n<p>If you looked at this dataset more closely, only 600 out of 1000 cases have radiologists unanimously agree about Covid presence in X-ray- this is insanely low agreement!</p>\n<p>On top of that, there are 150 cases where all radiologists think that there is no Covid (although tests have it confirmed)</p>",
      "rawMarkdown": "I have shared a very useful dataset, which contains 1000 X-rays, which were rated by 3 different radiologists using the same taxonomy (`\"negative\", \"typical\", \"indeterminate\", \"atypical\"`) as in the competition.\n\nhttps://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests\nPlease use it wisely!\n\nIf you looked at this dataset more closely, only 600 out of 1000 cases have radiologists unanimously agree about Covid presence in X-ray- this is insanely low agreement!\n\nOn top of that, there are 150 cases where all radiologists think that there is no Covid (although tests have it confirmed)\n\n\n\n",
      "votes": 111
    },
    {
      "id": 1379605,
      "postDate": "2021-07-07T13:35:28.750Z",
      "content": "<p><code>this dataset is part of official training/test set of the competition.</code><br>\nIs this part of the pictures of this competition，and you just had three different radiologists re-marked?<br>\nSo how do I correlate your data's StudyInstanceUID with the official StudyInstanceUID? They don't look the same.</p>",
      "rawMarkdown": "`this dataset is part of official training/test set of the competition.`\nIs this part of the pictures of this competition，and you just had three different radiologists re-marked?\nSo how do I correlate your data's StudyInstanceUID with the official StudyInstanceUID? They don't look the same.",
      "votes": 1
    },
    {
      "id": 1349102,
      "postDate": "2021-06-14T14:16:57.453Z",
      "content": "<p>Thanks for sharing the resource. Intriguing, the baseline model plateaus at ~65% accuracy. Arguably, we are learning the expert knowledge from <strong>the</strong> annotator in this competition.</p>",
      "rawMarkdown": "Thanks for sharing the resource. Intriguing, the baseline model plateaus at ~65% accuracy. Arguably, we are learning the expert knowledge from **the** annotator in this competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1367712,
          "postDate": "2021-06-28T03:44:58.370Z",
          "content": "<p><img src=\"https://i.ibb.co/xgh6zvn/agree.png\" alt=\"\"> <br>\n--Update--<br>\nDid a quick analysis. The unanimous agreement % is actually lower than 60%, if we exclude the images with only 1 annotator. Expert often reaches consensus on the <strong>no finding</strong> case, but disagree more frequently on other categories.  </p>\n<p>At the study level when multiple X-ray images were taken over the course of a disease, 41.8% received consistent label (e.g., always negative/indeterminate/typical/atypical across different time point), while 35.8% received 2 and 22.3% received 3 distinct labels… Here label means majority vote. This highlighted the uncertainty and descriptive nature of appearance definition.</p>",
          "rawMarkdown": "![](https://i.ibb.co/xgh6zvn/agree.png) \n--Update--\nDid a quick analysis. The unanimous agreement % is actually lower than 60%, if we exclude the images with only 1 annotator. Expert often reaches consensus on the **no finding** case, but disagree more frequently on other categories.  \n\nAt the study level when multiple X-ray images were taken over the course of a disease, 41.8% received consistent label (e.g., always negative/indeterminate/typical/atypical across different time point), while 35.8% received 2 and 22.3% received 3 distinct labels... Here label means majority vote. This highlighted the uncertainty and descriptive nature of appearance definition."
        }
      ]
    },
    {
      "id": 1314883,
      "postDate": "2021-05-19T12:33:01.427Z",
      "content": "<p>Is it possible that those 150 patients are at an early stage of infection / or asymptomatic? It's strange because my dad had a CORADS-5 and still his tests were Negative, which is totally the opposite</p>\n<p>Would be much helpful for my personal studies if you have which 150s are those. Thanks for this!</p>",
      "rawMarkdown": "Is it possible that those 150 patients are at an early stage of infection / or asymptomatic? It's strange because my dad had a CORADS-5 and still his tests were Negative, which is totally the opposite\n\nWould be much helpful for my personal studies if you have which 150s are those. Thanks for this!",
      "votes": 1,
      "replies": [
        {
          "id": 1314931,
          "postDate": "2021-05-19T13:10:02.823Z",
          "content": "<p>X-rays have low relevance for early Covid-19 detection and actually are not recommended by clinicians. Having these 150 cases in this dataset is a clear reason why :) False Negatives coming from X-ray reading are very bad compared to other types of Covid-19 testing.</p>\n<p>However, X-rays can be useful to track the progress of the treatment by comparing two X-rays and doing some quantitative analysis - that is why competition has this as a segmentation task.</p>\n<p>As for you dad case, any test has its own has False Negatives. It's just the tests have failed. You could actually try to lookup False Positive and False Negative rates for tests your dad took. You would be surprised about their accuracy :)</p>\n<p>As for those 150 cases - please lookup the metadata file in the dataset. These cases are labeled as <code>label = Negative, Negative, Negative</code></p>",
          "rawMarkdown": "X-rays have low relevance for early Covid-19 detection and actually are not recommended by clinicians. Having these 150 cases in this dataset is a clear reason why :) False Negatives coming from X-ray reading are very bad compared to other types of Covid-19 testing.\n\nHowever, X-rays can be useful to track the progress of the treatment by comparing two X-rays and doing some quantitative analysis - that is why competition has this as a segmentation task.\n\nAs for you dad case, any test has its own has False Negatives. It's just the tests have failed. You could actually try to lookup False Positive and False Negative rates for tests your dad took. You would be surprised about their accuracy :)\n\nAs for those 150 cases - please lookup the metadata file in the dataset. These cases are labeled as `label = Negative, Negative, Negative`",
          "votes": 3
        },
        {
          "id": 1315354,
          "postDate": "2021-05-19T18:18:09.183Z",
          "content": "<p>Make sense. Thanks!</p>",
          "rawMarkdown": "Make sense. Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1324213,
      "postDate": "2021-05-26T17:53:37.377Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. I think the dataset falls under the <code>CC BY-NC 4.0</code> license which states, <strong>\"You may not use the material for commercial purposes\"</strong>. I'm wondering if this will be problematic for this competition.</p>",
      "rawMarkdown": "Hi @raddar. I think the dataset falls under the `CC BY-NC 4.0` license which states, **\"You may not use the material for commercial purposes\"**. I'm wondering if this will be problematic for this competition.",
      "votes": 2,
      "replies": [
        {
          "id": 1324731,
          "postDate": "2021-05-27T07:36:30.613Z",
          "content": "<p>It should not be a problem as a competition is organized by SIIM which is a nonprofit organization</p>",
          "rawMarkdown": "It should not be a problem as a competition is organized by SIIM which is a nonprofit organization",
          "votes": 3
        },
        {
          "id": 1348869,
          "postDate": "2021-06-14T10:50:04.833Z",
          "content": "<p>I fear this is deep in the grey area as it says specifically that it prohibits any \"monetary compensation\" which is something that IS offered by the competition orginizers even if they are a non profit organization.</p>",
          "rawMarkdown": "I fear this is deep in the grey area as it says specifically that it prohibits any \"monetary compensation\" which is something that IS offered by the competition orginizers even if they are a non profit organization."
        },
        {
          "id": 1350195,
          "postDate": "2021-06-15T10:27:34.920Z",
          "content": "<p><a href=\"https://www.kaggle.com/omeram\" target=\"_blank\">@omeram</a> this dataset is part of official training/test set of the competition.</p>",
          "rawMarkdown": "@omeram this dataset is part of official training/test set of the competition."
        }
      ]
    },
    {
      "id": 1367723,
      "postDate": "2021-06-28T03:54:02.873Z",
      "content": "<p>Thanks for providing us the data. <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <br>\nI saw many rows with more than 3 labels when splitting using comma. You said that the data was annotated by 3 different rads then why do we have this issue? Or am I misunderstanding something ?</p>",
      "rawMarkdown": "Thanks for providing us the data. @raddar \nI saw many rows with more than 3 labels when splitting using comma. You said that the data was annotated by 3 different rads then why do we have this issue? Or am I misunderstanding something ?",
      "replies": [
        {
          "id": 1373208,
          "postDate": "2021-07-02T10:04:24.590Z",
          "content": "<p>Same question, also in most cases the radiologists disagree with each other 😂</p>",
          "rawMarkdown": "Same question, also in most cases the radiologists disagree with each other 😂"
        }
      ]
    },
    {
      "id": 1355573,
      "postDate": "2021-06-18T11:26:32.970Z",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> any plan on publishing the <strong>BIMCV</strong> dataset? The data is approx ~400GB.</p>",
      "rawMarkdown": "@raddar any plan on publishing the **BIMCV** dataset? The data is approx ~400GB.",
      "replies": [
        {
          "id": 1367598,
          "postDate": "2021-06-27T22:45:32.587Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1345119,
      "postDate": "2021-06-11T10:26:57.927Z",
      "content": "<p>Very insightful. It's certainly worth exploring. Thank you 👍</p>",
      "rawMarkdown": "Very insightful. It's certainly worth exploring. Thank you 👍"
    },
    {
      "id": 1325346,
      "postDate": "2021-05-27T17:12:32.607Z",
      "content": "<p>Around <strong>11</strong> rows contain <strong>NaN</strong> as labels. Any idea what do they mean? <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a></p>",
      "rawMarkdown": "Around **11** rows contain **NaN** as labels. Any idea what do they mean? @raddar",
      "replies": [
        {
          "id": 1326056,
          "postDate": "2021-05-28T07:34:12.033Z",
          "content": "<p>I don't know. drop these I suppose… </p>\n<p>By the way, does this data help with your score? impressive so far</p>",
          "rawMarkdown": "I don't know. drop these I suppose... \n\nBy the way, does this data help with your score? impressive so far",
          "votes": 1
        },
        {
          "id": 1326061,
          "postDate": "2021-05-28T07:37:27.713Z",
          "content": "<p>I haven't used it on my pipeline yet.</p>",
          "rawMarkdown": "I haven't used it on my pipeline yet."
        }
      ]
    },
    {
      "id": 1325248,
      "postDate": "2021-05-27T15:40:37.140Z",
      "content": "<p>Is external dataset allowed in this competition? In previous competitions I remember there were threads with list of allowed datasets, is this competition different?</p>",
      "rawMarkdown": "Is external dataset allowed in this competition? In previous competitions I remember there were threads with list of allowed datasets, is this competition different?",
      "replies": [
        {
          "id": 1325310,
          "postDate": "2021-05-27T16:33:45.870Z",
          "content": "<p>maybe kaggle organizers forgot to create a forum thread for that.</p>",
          "rawMarkdown": "maybe kaggle organizers forgot to create a forum thread for that."
        },
        {
          "id": 1325330,
          "postDate": "2021-05-27T16:51:29.493Z",
          "content": "<p>let's hope they will comment on that, otherwise it will be gray area</p>",
          "rawMarkdown": "let's hope they will comment on that, otherwise it will be gray area"
        },
        {
          "id": 1325334,
          "postDate": "2021-05-27T16:57:32.760Z",
          "content": "<p>why? rules are pretty clear about external data </p>",
          "rawMarkdown": "why? rules are pretty clear about external data ",
          "votes": 1
        },
        {
          "id": 1325390,
          "postDate": "2021-05-27T17:56:18.557Z",
          "content": "<p>I was thinking this thread with external data is needed for each competition</p>",
          "rawMarkdown": "I was thinking this thread with external data is needed for each competition"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1379605,
      "author_name": "AndyX",
      "author_url": "",
      "post_date": "2021-07-07T13:35:28.750000",
      "content": "<p><code>this dataset is part of official training/test set of the competition.</code><br>\nIs this part of the pictures of this competition，and you just had three different radiologists re-marked?<br>\nSo how do I correlate your data's StudyInstanceUID with the official StudyInstanceUID? They don't look the same.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1349102,
      "author_name": "human intelligence",
      "author_url": "",
      "post_date": "2021-06-14T14:16:57.453000",
      "content": "<p>Thanks for sharing the resource. Intriguing, the baseline model plateaus at ~65% accuracy. Arguably, we are learning the expert knowledge from <strong>the</strong> annotator in this competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1367712,
          "author_name": "human intelligence",
          "author_url": "",
          "post_date": "2021-06-28T03:44:58.370000",
          "content": "<p><img src=\"https://i.ibb.co/xgh6zvn/agree.png\" alt=\"\"> <br>\n--Update--<br>\nDid a quick analysis. The unanimous agreement % is actually lower than 60%, if we exclude the images with only 1 annotator. Expert often reaches consensus on the <strong>no finding</strong> case, but disagree more frequently on other categories.  </p>\n<p>At the study level when multiple X-ray images were taken over the course of a disease, 41.8% received consistent label (e.g., always negative/indeterminate/typical/atypical across different time point), while 35.8% received 2 and 22.3% received 3 distinct labels… Here label means majority vote. This highlighted the uncertainty and descriptive nature of appearance definition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1314883,
      "author_name": "Shahebaz Mohammad",
      "author_url": "",
      "post_date": "2021-05-19T12:33:01.427000",
      "content": "<p>Is it possible that those 150 patients are at an early stage of infection / or asymptomatic? It's strange because my dad had a CORADS-5 and still his tests were Negative, which is totally the opposite</p>\n<p>Would be much helpful for my personal studies if you have which 150s are those. Thanks for this!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1314931,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-05-19T13:10:02.823000",
          "content": "<p>X-rays have low relevance for early Covid-19 detection and actually are not recommended by clinicians. Having these 150 cases in this dataset is a clear reason why :) False Negatives coming from X-ray reading are very bad compared to other types of Covid-19 testing.</p>\n<p>However, X-rays can be useful to track the progress of the treatment by comparing two X-rays and doing some quantitative analysis - that is why competition has this as a segmentation task.</p>\n<p>As for you dad case, any test has its own has False Negatives. It's just the tests have failed. You could actually try to lookup False Positive and False Negative rates for tests your dad took. You would be surprised about their accuracy :)</p>\n<p>As for those 150 cases - please lookup the metadata file in the dataset. These cases are labeled as <code>label = Negative, Negative, Negative</code></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1315354,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2021-05-19T18:18:09.183000",
          "content": "<p>Make sense. Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1324213,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-05-26T17:53:37.377000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. I think the dataset falls under the <code>CC BY-NC 4.0</code> license which states, <strong>\"You may not use the material for commercial purposes\"</strong>. I'm wondering if this will be problematic for this competition.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1324731,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-05-27T07:36:30.613000",
          "content": "<p>It should not be a problem as a competition is organized by SIIM which is a nonprofit organization</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1348869,
          "author_name": "Omer A.",
          "author_url": "",
          "post_date": "2021-06-14T10:50:04.833000",
          "content": "<p>I fear this is deep in the grey area as it says specifically that it prohibits any \"monetary compensation\" which is something that IS offered by the competition orginizers even if they are a non profit organization.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1350195,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-06-15T10:27:34.920000",
          "content": "<p><a href=\"https://www.kaggle.com/omeram\" target=\"_blank\">@omeram</a> this dataset is part of official training/test set of the competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1367723,
      "author_name": "Liam Nguyen",
      "author_url": "",
      "post_date": "2021-06-28T03:54:02.873000",
      "content": "<p>Thanks for providing us the data. <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <br>\nI saw many rows with more than 3 labels when splitting using comma. You said that the data was annotated by 3 different rads then why do we have this issue? Or am I misunderstanding something ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1373208,
          "author_name": "Logic",
          "author_url": "",
          "post_date": "2021-07-02T10:04:24.590000",
          "content": "<p>Same question, also in most cases the radiologists disagree with each other 😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1355573,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-06-18T11:26:32.970000",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> any plan on publishing the <strong>BIMCV</strong> dataset? The data is approx ~400GB.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1367598,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-06-27T22:45:32.587000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1345119,
      "author_name": "README",
      "author_url": "",
      "post_date": "2021-06-11T10:26:57.927000",
      "content": "<p>Very insightful. It's certainly worth exploring. Thank you 👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1325346,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-05-27T17:12:32.607000",
      "content": "<p>Around <strong>11</strong> rows contain <strong>NaN</strong> as labels. Any idea what do they mean? <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1326056,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-05-28T07:34:12.033000",
          "content": "<p>I don't know. drop these I suppose… </p>\n<p>By the way, does this data help with your score? impressive so far</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1326061,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-05-28T07:37:27.713000",
          "content": "<p>I haven't used it on my pipeline yet.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1325248,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2021-05-27T15:40:37.140000",
      "content": "<p>Is external dataset allowed in this competition? In previous competitions I remember there were threads with list of allowed datasets, is this competition different?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1325310,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-05-27T16:33:45.870000",
          "content": "<p>maybe kaggle organizers forgot to create a forum thread for that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325330,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-27T16:51:29.493000",
          "content": "<p>let's hope they will comment on that, otherwise it will be gray area</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325334,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-05-27T16:57:32.760000",
          "content": "<p>why? rules are pretty clear about external data </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1325390,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-27T17:56:18.557000",
          "content": "<p>I was thinking this thread with external data is needed for each competition</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1313908": "I have shared a very useful dataset, which contains 1000 X-rays, which were rated by 3 different radiologists using the same taxonomy (`\"negative\", \"typical\", \"indeterminate\", \"atypical\"`) as in the competition.\n\nhttps://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests\nPlease use it wisely!\n\nIf you looked at this dataset more closely, only 600 out of 1000 cases have radiologists unanimously agree about Covid presence in X-ray- this is insanely low agreement!\n\nOn top of that, there are 150 cases where all radiologists think that there is no Covid (although tests have it confirmed)\n\n\n\n",
    "1379605": "`this dataset is part of official training/test set of the competition.`\nIs this part of the pictures of this competition，and you just had three different radiologists re-marked?\nSo how do I correlate your data's StudyInstanceUID with the official StudyInstanceUID? They don't look the same.",
    "1349102": "Thanks for sharing the resource. Intriguing, the baseline model plateaus at ~65% accuracy. Arguably, we are learning the expert knowledge from **the** annotator in this competition.",
    "1314883": "Is it possible that those 150 patients are at an early stage of infection / or asymptomatic? It's strange because my dad had a CORADS-5 and still his tests were Negative, which is totally the opposite\n\nWould be much helpful for my personal studies if you have which 150s are those. Thanks for this!",
    "1324213": "Hi @raddar. I think the dataset falls under the `CC BY-NC 4.0` license which states, **\"You may not use the material for commercial purposes\"**. I'm wondering if this will be problematic for this competition.",
    "1367723": "Thanks for providing us the data. @raddar \nI saw many rows with more than 3 labels when splitting using comma. You said that the data was annotated by 3 different rads then why do we have this issue? Or am I misunderstanding something ?",
    "1355573": "@raddar any plan on publishing the **BIMCV** dataset? The data is approx ~400GB.",
    "1345119": "Very insightful. It's certainly worth exploring. Thank you 👍",
    "1325346": "Around **11** rows contain **NaN** as labels. Any idea what do they mean? @raddar",
    "1325248": "Is external dataset allowed in this competition? In previous competitions I remember there were threads with list of allowed datasets, is this competition different?"
  }
}