{
  "id": 220927,
  "title": "kaggle Competition DevOps : Model -> Data -> Bias",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/220927",
  "author_name": "hengck23",
  "post_date": "2021-02-20T05:54:36.833000",
  "votes": 50,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have just completed RFCx frog/bird species audio detection and back to RANZCR challenge here.<br>\nI want to introduce a development framework that I find useful for Kaggle challenge.</p>\n<p>It has 3 steps:</p>\n<ol>\n<li><p>modeling  </p>\n<ul>\n<li>problem setup, algorithm design</li></ul></li>\n<li><p>data  </p>\n<ul>\n<li>scaling up with more data</li></ul></li>\n<li><p>bias</p>\n<ul>\n<li>reduce shakeup by reducing test/train data bias</li></ul></li>\n</ol>\n<p>You can also treat them as tasks for <br>\ndata scientist --&gt; data engineer --&gt; data magician, or<br>\nexploration --&gt; exploitation --&gt; explosion</p>\n<p>examples are (search for my comment in the post below) : <br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339</a><br>\n<a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189344\" target=\"_blank\">https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189344</a></p>\n<hr>\n<p>How to apply this in this RANZCR challenge?</p>\n<ol>\n<li><p>modeling  </p>\n<ul>\n<li><p>baseline framework is CNN image classifier. Can we do better than this? The key is if you can find use of the extra annotation or not. There are already some discussions on it, e.g. using the teacher distillation approach, as auxiliary segmentation loss, object detection approach (detect tube endpoint), etc …</p></li>\n<li><p>other work includes finding the best image size, network architecture, training hyperparameters etc …. or other new information on the data</p></li></ul></li>\n\n<li><p>data  </p>\n<ul>\n<li><p>only part of the train data is annotated with extra information. If you manage to find good use of this extra annotation, your model may improve if you can extend this annotation in the other un-annotated train data. E.g. hand label, semi-supervised (like the pseudo label). You can also think of other new self-supervised methods as well.</p></li>\n<li><p>If you can not extend this annotation, why don't create them? GAN, synthetic data, novel augmentations, etc</p></li>\n<li><p>The key is to use more data to amplify your modeling performance above</p></li></ul></li>\n<li><p>bais</p>\n<ul>\n<li><p>traditionally, this has been called the magic in kaggle. An example is: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a></p></li>\n<li><p>Another example is to add bias to the kaggle regression challenges, (e.g. in some forecast challenge, just adding 0.1 to your predicted results improve LB scores) </p></li>\n<li><p>I haven't probe and check the data for this competition yet. maybe I will comment more on this later.</p></li></ul></li>\n</ol>",
  "messages": [
    {
      "id": 1211297,
      "postDate": "2021-02-20T05:54:36.833Z",
      "content": "<p>I have just completed RFCx frog/bird species audio detection and back to RANZCR challenge here.<br>\nI want to introduce a development framework that I find useful for Kaggle challenge.</p>\n<p>It has 3 steps:</p>\n<ol>\n<li><p>modeling  </p>\n<ul>\n<li>problem setup, algorithm design</li></ul></li>\n<li><p>data  </p>\n<ul>\n<li>scaling up with more data</li></ul></li>\n<li><p>bias</p>\n<ul>\n<li>reduce shakeup by reducing test/train data bias</li></ul></li>\n</ol>\n<p>You can also treat them as tasks for <br>\ndata scientist --&gt; data engineer --&gt; data magician, or<br>\nexploration --&gt; exploitation --&gt; explosion</p>\n<p>examples are (search for my comment in the post below) : <br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339</a><br>\n<a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189344\" target=\"_blank\">https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189344</a></p>\n<hr>\n<p>How to apply this in this RANZCR challenge?</p>\n<ol>\n<li><p>modeling  </p>\n<ul>\n<li><p>baseline framework is CNN image classifier. Can we do better than this? The key is if you can find use of the extra annotation or not. There are already some discussions on it, e.g. using the teacher distillation approach, as auxiliary segmentation loss, object detection approach (detect tube endpoint), etc …</p></li>\n<li><p>other work includes finding the best image size, network architecture, training hyperparameters etc …. or other new information on the data</p></li></ul></li>\n\n<li><p>data  </p>\n<ul>\n<li><p>only part of the train data is annotated with extra information. If you manage to find good use of this extra annotation, your model may improve if you can extend this annotation in the other un-annotated train data. E.g. hand label, semi-supervised (like the pseudo label). You can also think of other new self-supervised methods as well.</p></li>\n<li><p>If you can not extend this annotation, why don't create them? GAN, synthetic data, novel augmentations, etc</p></li>\n<li><p>The key is to use more data to amplify your modeling performance above</p></li></ul></li>\n<li><p>bais</p>\n<ul>\n<li><p>traditionally, this has been called the magic in kaggle. An example is: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a></p></li>\n<li><p>Another example is to add bias to the kaggle regression challenges, (e.g. in some forecast challenge, just adding 0.1 to your predicted results improve LB scores) </p></li>\n<li><p>I haven't probe and check the data for this competition yet. maybe I will comment more on this later.</p></li></ul></li>\n</ol>",
      "rawMarkdown": "I have just completed RFCx frog/bird species audio detection and back to RANZCR challenge here.\nI want to introduce a development framework that I find useful for Kaggle challenge.\n\nIt has 3 steps:\n1. modeling  \n  - problem setup, algorithm design\n\n2. data  \n  - scaling up with more data\n\n3. bias\n  - reduce shakeup by reducing test/train data bias\n\nYou can also treat them as tasks for \ndata scientist --> data engineer --> data magician, or\nexploration --> exploitation --> explosion\n\nexamples are (search for my comment in the post below) : \nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339\nhttps://www.kaggle.com/c/stanford-covid-vaccine/discussion/189344\n\n---\n\nHow to apply this in this RANZCR challenge?\n\n\n1. modeling  \n\n  - baseline framework is CNN image classifier. Can we do better than this? The key is if you can find use of the extra annotation or not. There are already some discussions on it, e.g. using the teacher distillation approach, as auxiliary segmentation loss, object detection approach (detect tube endpoint), etc ...\n\n  - other work includes finding the best image size, network architecture, training hyperparameters etc .... or other new information on the data\n\n\n\n2. data  \n  - only part of the train data is annotated with extra information. If you manage to find good use of this extra annotation, your model may improve if you can extend this annotation in the other un-annotated train data. E.g. hand label, semi-supervised (like the pseudo label). You can also think of other new self-supervised methods as well.\n\n  - If you can not extend this annotation, why don't create them? GAN, synthetic data, novel augmentations, etc\n\n  - The key is to use more data to amplify your modeling performance above\n\n\n3. bais\n  - traditionally, this has been called the magic in kaggle. An example is: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\n\n  - Another example is to add bias to the kaggle regression challenges, (e.g. in some forecast challenge, just adding 0.1 to your predicted results improve LB scores) \n\n  - I haven't probe and check the data for this competition yet. maybe I will comment more on this later.",
      "votes": 50
    },
    {
      "id": 1214940,
      "postDate": "2021-02-23T08:04:08.557Z",
      "content": "<p>How would bias change a ROC_AUC score?</p>",
      "rawMarkdown": "How would bias change a ROC_AUC score?",
      "votes": 3
    },
    {
      "id": 1227532,
      "postDate": "2021-03-05T16:17:07.567Z",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "rawMarkdown": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
      "votes": 1
    },
    {
      "id": 1224339,
      "postDate": "2021-03-02T16:38:55.093Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for sharing this idea. Can you share any kernel which uses semi-supervised for this competition? Thanks in advance.</p>",
      "rawMarkdown": "@hengck23 Thanks for sharing this idea. Can you share any kernel which uses semi-supervised for this competition? Thanks in advance."
    },
    {
      "id": 1214935,
      "postDate": "2021-02-23T08:01:04.333Z",
      "content": "<p>Thanks for sharing, very useful information!</p>",
      "rawMarkdown": "Thanks for sharing, very useful information!"
    },
    {
      "id": 1227498,
      "postDate": "2021-03-05T15:36:42.130Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1217949,
      "postDate": "2021-02-25T12:38:03.803Z",
      "content": "<p>Thanks for sharing,</p>",
      "rawMarkdown": "Thanks for sharing,",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1214940,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2021-02-23T08:04:08.557000",
      "content": "<p>How would bias change a ROC_AUC score?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1227532,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-03-05T16:17:07.567000",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1224339,
      "author_name": "Md. Masud Rana",
      "author_url": "",
      "post_date": "2021-03-02T16:38:55.093000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for sharing this idea. Can you share any kernel which uses semi-supervised for this competition? Thanks in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1214935,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-02-23T08:01:04.333000",
      "content": "<p>Thanks for sharing, very useful information!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1227498,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-05T15:36:42.130000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1217949,
      "author_name": "Patrick",
      "author_url": "",
      "post_date": "2021-02-25T12:38:03.803000",
      "content": "<p>Thanks for sharing,</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1211297": "I have just completed RFCx frog/bird species audio detection and back to RANZCR challenge here.\nI want to introduce a development framework that I find useful for Kaggle challenge.\n\nIt has 3 steps:\n1. modeling  \n  - problem setup, algorithm design\n\n2. data  \n  - scaling up with more data\n\n3. bias\n  - reduce shakeup by reducing test/train data bias\n\nYou can also treat them as tasks for \ndata scientist --> data engineer --> data magician, or\nexploration --> exploitation --> explosion\n\nexamples are (search for my comment in the post below) : \nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339\nhttps://www.kaggle.com/c/stanford-covid-vaccine/discussion/189344\n\n---\n\nHow to apply this in this RANZCR challenge?\n\n\n1. modeling  \n\n  - baseline framework is CNN image classifier. Can we do better than this? The key is if you can find use of the extra annotation or not. There are already some discussions on it, e.g. using the teacher distillation approach, as auxiliary segmentation loss, object detection approach (detect tube endpoint), etc ...\n\n  - other work includes finding the best image size, network architecture, training hyperparameters etc .... or other new information on the data\n\n\n\n2. data  \n  - only part of the train data is annotated with extra information. If you manage to find good use of this extra annotation, your model may improve if you can extend this annotation in the other un-annotated train data. E.g. hand label, semi-supervised (like the pseudo label). You can also think of other new self-supervised methods as well.\n\n  - If you can not extend this annotation, why don't create them? GAN, synthetic data, novel augmentations, etc\n\n  - The key is to use more data to amplify your modeling performance above\n\n\n3. bais\n  - traditionally, this has been called the magic in kaggle. An example is: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\n\n  - Another example is to add bias to the kaggle regression challenges, (e.g. in some forecast challenge, just adding 0.1 to your predicted results improve LB scores) \n\n  - I haven't probe and check the data for this competition yet. maybe I will comment more on this later.",
    "1214940": "How would bias change a ROC_AUC score?",
    "1227532": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
    "1224339": "@hengck23 Thanks for sharing this idea. Can you share any kernel which uses semi-supervised for this competition? Thanks in advance.",
    "1214935": "Thanks for sharing, very useful information!",
    "1227498": "",
    "1217949": "Thanks for sharing,"
  }
}