{
  "id": 30395,
  "title": "Weird trend with Additional Data",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30395",
  "author_name": "",
  "post_date": "2017-03-20T08:50:07.907403700Z",
  "votes": 5,
  "comment_count": 5,
  "views": 1,
  "content": "<p>Hi everyone,</p>\n\n<p>I wanted to ask if anyone else notices weird behavior whilst training on the additional data.</p>\n\n<p>The model which gave me my best result on the leaderboard was trained using images from the \"train\" data only, using 10% holdout for validation.</p>\n\n<p>When I train the same model using images from the \"additional\" data set, and the images from \"train\" data set as validation, <strong>the model just doesn't train. validation loss never improves, only oscillates between few SPECIFIC values.</strong> </p>\n\n<p>I am sure it isn't a problem with the learning rate, because I have used both NAG and Adam optimizers. </p>\n\n<p>Images from both \"train\" and \"additional\" are preprocessed in the exact same way.</p>\n\n<p>Has anyone experienced something similar?</p>",
  "messages": [
    {
      "id": "169244",
      "postDate": "03/20/2017 08:50:07",
      "content": "<p>Hi everyone,</p>\n\n<p>I wanted to ask if anyone else notices weird behavior whilst training on the additional data.</p>\n\n<p>The model which gave me my best result on the leaderboard was trained using images from the \"train\" data only, using 10% holdout for validation.</p>\n\n<p>When I train the same model using images from the \"additional\" data set, and the images from \"train\" data set as validation, <strong>the model just doesn't train. validation loss never improves, only oscillates between few SPECIFIC values.</strong> </p>\n\n<p>I am sure it isn't a problem with the learning rate, because I have used both NAG and Adam optimizers. </p>\n\n<p>Images from both \"train\" and \"additional\" are preprocessed in the exact same way.</p>\n\n<p>Has anyone experienced something similar?</p>",
      "rawMarkdown": "Hi everyone,\n\nI wanted to ask if anyone else notices weird behavior whilst training on the additional data.\n\nThe model which gave me my best result on the leaderboard was trained using images from the \"train\" data only, using 10% holdout for validation.\n\nWhen I train the same model using images from the \"additional\" data set, and the images from \"train\" data set as validation, **the model just doesn't train. validation loss never improves, only oscillates between few SPECIFIC values.** \n\nI am sure it isn't a problem with the learning rate, because I have used both NAG and Adam optimizers. \n\nImages from both \"train\" and \"additional\" are preprocessed in the exact same way.\n\nHas anyone experienced something similar?",
      "votes": null
    },
    {
      "id": "169317",
      "postDate": "03/20/2017 14:46:57",
      "content": "<p>bump</p>",
      "rawMarkdown": "bump",
      "votes": null
    },
    {
      "id": "169529",
      "postDate": "03/21/2017 10:40:52",
      "content": "<p>Hey, I am facing the same problem, how did you resolve the issue?</p>",
      "rawMarkdown": "Hey, I am facing the same problem, how did you resolve the issue?",
      "votes": null
    },
    {
      "id": "169536",
      "postDate": "03/21/2017 11:58:58",
      "content": "<p>I haven't actually. Training using the Additional set is still bugging me.</p>",
      "rawMarkdown": "I haven't actually. Training using the Additional set is still bugging me.",
      "votes": null
    },
    {
      "id": "169638",
      "postDate": "03/21/2017 21:14:11",
      "content": "<p>Hm, \"ugly\" images are the norm there. 1 in 5 is heavily blurred, almost all of them have some pathological conditions. They would have to be manually curated somehow. Oh, wait! They were :( </p>\n\n<p>And another thing. Afaik in additional.zip there are more than 1 picture per patient. From my experience is easier for a classifier to pick up the patient rather than the disease. Training on additional would make a good patient detector.  Chances are that only few patients from the train are in additional. And there is exactly 1 picture per patient.  Please tell me if I remember wrong!</p>",
      "rawMarkdown": "Hm, \"ugly\" images are the norm there. 1 in 5 is heavily blurred, almost all of them have some pathological conditions. They would have to be manually curated somehow. Oh, wait! They were :( \n\nAnd another thing. Afaik in additional.zip there are more than 1 picture per patient. From my experience is easier for a classifier to pick up the patient rather than the disease. Training on additional would make a good patient detector.  Chances are that only few patients from the train are in additional. And there is exactly 1 picture per patient.  Please tell me if I remember wrong!",
      "votes": null
    },
    {
      "id": "169691",
      "postDate": "03/22/2017 05:25:46",
      "content": "<p>I second that. I spent a few hours last night manually going through each of the images in the additional set (didn't make it past Type_1 though) and some images were downright useless!</p>",
      "rawMarkdown": "I second that. I spent a few hours last night manually going through each of the images in the additional set (didn't make it past Type_1 though) and some images were downright useless!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 169317,
      "author_name": "yadavsarthak",
      "author_url": "",
      "post_date": "03/20/2017 14:46:57",
      "content": "<p>bump</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 169529,
      "author_name": "prajwalkr",
      "author_url": "",
      "post_date": "03/21/2017 10:40:52",
      "content": "<p>Hey, I am facing the same problem, how did you resolve the issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 169536,
          "author_name": "yadavsarthak",
          "author_url": "",
          "post_date": "03/21/2017 11:58:58",
          "content": "<p>I haven't actually. Training using the Additional set is still bugging me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 169638,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "03/21/2017 21:14:11",
      "content": "<p>Hm, \"ugly\" images are the norm there. 1 in 5 is heavily blurred, almost all of them have some pathological conditions. They would have to be manually curated somehow. Oh, wait! They were :( </p>\n\n<p>And another thing. Afaik in additional.zip there are more than 1 picture per patient. From my experience is easier for a classifier to pick up the patient rather than the disease. Training on additional would make a good patient detector.  Chances are that only few patients from the train are in additional. And there is exactly 1 picture per patient.  Please tell me if I remember wrong!</p>",
      "votes": null,
      "replies": [
        {
          "id": 169691,
          "author_name": "yadavsarthak",
          "author_url": "",
          "post_date": "03/22/2017 05:25:46",
          "content": "<p>I second that. I spent a few hours last night manually going through each of the images in the additional set (didn't make it past Type_1 though) and some images were downright useless!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "169244": "Hi everyone,\n\nI wanted to ask if anyone else notices weird behavior whilst training on the additional data.\n\nThe model which gave me my best result on the leaderboard was trained using images from the \"train\" data only, using 10% holdout for validation.\n\nWhen I train the same model using images from the \"additional\" data set, and the images from \"train\" data set as validation, **the model just doesn't train. validation loss never improves, only oscillates between few SPECIFIC values.** \n\nI am sure it isn't a problem with the learning rate, because I have used both NAG and Adam optimizers. \n\nImages from both \"train\" and \"additional\" are preprocessed in the exact same way.\n\nHas anyone experienced something similar?",
    "169317": "bump",
    "169529": "Hey, I am facing the same problem, how did you resolve the issue?",
    "169536": "I haven't actually. Training using the Additional set is still bugging me.",
    "169638": "Hm, \"ugly\" images are the norm there. 1 in 5 is heavily blurred, almost all of them have some pathological conditions. They would have to be manually curated somehow. Oh, wait! They were :( \n\nAnd another thing. Afaik in additional.zip there are more than 1 picture per patient. From my experience is easier for a classifier to pick up the patient rather than the disease. Training on additional would make a good patient detector.  Chances are that only few patients from the train are in additional. And there is exactly 1 picture per patient.  Please tell me if I remember wrong!",
    "169691": "I second that. I spent a few hours last night manually going through each of the images in the additional set (didn't make it past Type_1 though) and some images were downright useless!"
  },
  "source": "meta"
}