{
  "id": 29493,
  "title": "Small vehicles training strategy",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/29493",
  "author_name": "",
  "post_date": "2017-03-02T12:15:04.187317Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, <br>\nI'm struggling to find a good strategy for small vehicles. <br>\nI have trained with cross-validation and n_folds 2 and 3.  I get test scores between 0.2 and 0.3, and when I submit the predictions the scores do not correlate and are between 0.02 and 0.06.</p>\n\n<p>I can't understand why my test scores do not correlate with the leaderboard. Could it be that the data in train is very different from the test data?</p>\n\n<p>Thanks <br>\nironbar</p>",
  "messages": [
    {
      "id": "164802",
      "postDate": "03/02/2017 12:15:04",
      "content": "<p>Hi, <br>\nI'm struggling to find a good strategy for small vehicles. <br>\nI have trained with cross-validation and n_folds 2 and 3.  I get test scores between 0.2 and 0.3, and when I submit the predictions the scores do not correlate and are between 0.02 and 0.06.</p>\n\n<p>I can't understand why my test scores do not correlate with the leaderboard. Could it be that the data in train is very different from the test data?</p>\n\n<p>Thanks <br>\nironbar</p>",
      "rawMarkdown": "Hi,   \nI'm struggling to find a good strategy for small vehicles.   \nI have trained with cross-validation and n_folds 2 and 3.  I get test scores between 0.2 and 0.3, and when I submit the predictions the scores do not correlate and are between 0.02 and 0.06.\n\nI can't understand why my test scores do not correlate with the leaderboard. Could it be that the data in train is very different from the test data?\n\nThanks   \nironbar",
      "votes": null
    },
    {
      "id": "164854",
      "postDate": "03/02/2017 15:48:03",
      "content": "<p>I know that feel bro. I think there are two reasons for low score:</p>\n\n<ul>\n<li>public test very different from trainset</li>\n<li>small surface area of all vehicles. so false positive detection may affect on metric dramatically</li>\n</ul>",
      "rawMarkdown": "I know that feel bro. I think there are two reasons for low score:\n\n - public test very different from trainset\n - small surface area of all vehicles. so false positive detection may affect on metric dramatically",
      "votes": null
    },
    {
      "id": "164855",
      "postDate": "03/02/2017 15:51:10",
      "content": "<p>We all know your pain. Those cars are driving me insane. I don't know what exactly method DSTL used to find and segment cars but it looks like nearly impossible to reproduce given the extremely low amount of training data.</p>",
      "rawMarkdown": "We all know your pain. Those cars are driving me insane. I don't know what exactly method DSTL used to find and segment cars but it looks like nearly impossible to reproduce given the extremely low amount of training data.",
      "votes": null
    },
    {
      "id": "164859",
      "postDate": "03/02/2017 16:01:26",
      "content": "<p>You know, just recently I submitted a prediction that was getting pretty decent scores on validation, but it got 0.00000 on the LB. You're not alone. My bet is that the main reason is that there are very very few cars on the public LB.</p>",
      "rawMarkdown": "You know, just recently I submitted a prediction that was getting pretty decent scores on validation, but it got 0.00000 on the LB. You're not alone. My bet is that the main reason is that there are very very few cars on the public LB.",
      "votes": null
    },
    {
      "id": "164876",
      "postDate": "03/02/2017 17:37:05",
      "content": "<p>Maybe most of the cars are on the private LB...</p>",
      "rawMarkdown": "Maybe most of the cars are on the private LB...",
      "votes": null
    },
    {
      "id": "164955",
      "postDate": "03/03/2017 01:53:10",
      "content": "<p>test score 0.2~0.3 is much better than mine, difference between training and test set is really a challenge in this game.</p>",
      "rawMarkdown": "test score 0.2~0.3 is much better than mine, difference between training and test set is really a challenge in this game.",
      "votes": null
    },
    {
      "id": "165044",
      "postDate": "03/03/2017 12:09:49",
      "content": "<p>In the following link they recomend to use a lot of different simple models in an ensemble when there's no correlation between cross-validation score and LB. <br>\n<a href=\"http://blog.kaggle.com/2017/02/06/seizure-prediction-competition-first-place-winners-interview-team-not-so-random-anymore-andriy-alexandre-feng-gilberto/?utm_source=Mailing+list&amp;utm_campaign=1283010c27-Kaggle_Newsletter_03-01-2017&amp;utm_medium=email&amp;utm_term=0_f42f9df1e1-1283010c27-399697965\">http://blog.kaggle.com/2017/02/06/seizure-prediction-competition-first-place-winners-interview-team-not-so-random-anymore-andriy-alexandre-feng-gilberto/?utm_source=Mailing+list&amp;utm_campaign=1283010c27-Kaggle_Newsletter_03-01-2017&amp;utm_medium=email&amp;utm_term=0_f42f9df1e1-1283010c27-399697965</a></p>",
      "rawMarkdown": "In the following link they recomend to use a lot of different simple models in an ensemble when there's no correlation between cross-validation score and LB.  \nhttp://blog.kaggle.com/2017/02/06/seizure-prediction-competition-first-place-winners-interview-team-not-so-random-anymore-andriy-alexandre-feng-gilberto/?utm_source=Mailing+list&utm_campaign=1283010c27-Kaggle_Newsletter_03-01-2017&utm_medium=email&utm_term=0_f42f9df1e1-1283010c27-399697965",
      "votes": null
    },
    {
      "id": "165057",
      "postDate": "03/03/2017 13:36:07",
      "content": "<p>I feel that making predictions for class 9 and 10 are very difficult.<br>\nOn the public LB, I get 0.00195 (class 9) and 0.00000 (class 10).</p>",
      "rawMarkdown": "I feel that making predictions for class 9 and 10 are very difficult.<br>\nOn the public LB, I get 0.00195 (class 9) and 0.00000 (class 10).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 164854,
      "author_name": "drn01z3",
      "author_url": "",
      "post_date": "03/02/2017 15:48:03",
      "content": "<p>I know that feel bro. I think there are two reasons for low score:</p>\n\n<ul>\n<li>public test very different from trainset</li>\n<li>small surface area of all vehicles. so false positive detection may affect on metric dramatically</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 164855,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "03/02/2017 15:51:10",
      "content": "<p>We all know your pain. Those cars are driving me insane. I don't know what exactly method DSTL used to find and segment cars but it looks like nearly impossible to reproduce given the extremely low amount of training data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 164859,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "03/02/2017 16:01:26",
      "content": "<p>You know, just recently I submitted a prediction that was getting pretty decent scores on validation, but it got 0.00000 on the LB. You're not alone. My bet is that the main reason is that there are very very few cars on the public LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 164876,
          "author_name": "ironbar",
          "author_url": "",
          "post_date": "03/02/2017 17:37:05",
          "content": "<p>Maybe most of the cars are on the private LB...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 164955,
      "author_name": "zeliek",
      "author_url": "",
      "post_date": "03/03/2017 01:53:10",
      "content": "<p>test score 0.2~0.3 is much better than mine, difference between training and test set is really a challenge in this game.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 165044,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "03/03/2017 12:09:49",
      "content": "<p>In the following link they recomend to use a lot of different simple models in an ensemble when there's no correlation between cross-validation score and LB. <br>\n<a href=\"http://blog.kaggle.com/2017/02/06/seizure-prediction-competition-first-place-winners-interview-team-not-so-random-anymore-andriy-alexandre-feng-gilberto/?utm_source=Mailing+list&amp;utm_campaign=1283010c27-Kaggle_Newsletter_03-01-2017&amp;utm_medium=email&amp;utm_term=0_f42f9df1e1-1283010c27-399697965\">http://blog.kaggle.com/2017/02/06/seizure-prediction-competition-first-place-winners-interview-team-not-so-random-anymore-andriy-alexandre-feng-gilberto/?utm_source=Mailing+list&amp;utm_campaign=1283010c27-Kaggle_Newsletter_03-01-2017&amp;utm_medium=email&amp;utm_term=0_f42f9df1e1-1283010c27-399697965</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 165057,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "03/03/2017 13:36:07",
      "content": "<p>I feel that making predictions for class 9 and 10 are very difficult.<br>\nOn the public LB, I get 0.00195 (class 9) and 0.00000 (class 10).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "164802": "Hi,   \nI'm struggling to find a good strategy for small vehicles.   \nI have trained with cross-validation and n_folds 2 and 3.  I get test scores between 0.2 and 0.3, and when I submit the predictions the scores do not correlate and are between 0.02 and 0.06.\n\nI can't understand why my test scores do not correlate with the leaderboard. Could it be that the data in train is very different from the test data?\n\nThanks   \nironbar",
    "164854": "I know that feel bro. I think there are two reasons for low score:\n\n - public test very different from trainset\n - small surface area of all vehicles. so false positive detection may affect on metric dramatically",
    "164855": "We all know your pain. Those cars are driving me insane. I don't know what exactly method DSTL used to find and segment cars but it looks like nearly impossible to reproduce given the extremely low amount of training data.",
    "164859": "You know, just recently I submitted a prediction that was getting pretty decent scores on validation, but it got 0.00000 on the LB. You're not alone. My bet is that the main reason is that there are very very few cars on the public LB.",
    "164876": "Maybe most of the cars are on the private LB...",
    "164955": "test score 0.2~0.3 is much better than mine, difference between training and test set is really a challenge in this game.",
    "165044": "In the following link they recomend to use a lot of different simple models in an ensemble when there's no correlation between cross-validation score and LB.  \nhttp://blog.kaggle.com/2017/02/06/seizure-prediction-competition-first-place-winners-interview-team-not-so-random-anymore-andriy-alexandre-feng-gilberto/?utm_source=Mailing+list&utm_campaign=1283010c27-Kaggle_Newsletter_03-01-2017&utm_medium=email&utm_term=0_f42f9df1e1-1283010c27-399697965",
    "165057": "I feel that making predictions for class 9 and 10 are very difficult.<br>\nOn the public LB, I get 0.00195 (class 9) and 0.00000 (class 10)."
  },
  "source": "meta"
}