{
  "id": 77286,
  "title": "How to create validation data for simaese network",
  "url": "/competitions/humpback-whale-identification/discussion/77286",
  "author_name": "",
  "post_date": "2019-01-11T05:43:28.088419200Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>How to create validation data for simaese network? Is there any work related?</p>",
  "messages": [
    {
      "id": "454090",
      "postDate": "01/11/2019 05:43:28",
      "content": "<p>How to create validation data for simaese network? Is there any work related?</p>",
      "rawMarkdown": "How to create validation data for simaese network? Is there any work related?",
      "votes": null
    },
    {
      "id": "454704",
      "postDate": "01/12/2019 03:39:43",
      "content": "<p>I am also puzzled by this question.</p>",
      "rawMarkdown": "I am also puzzled by this question.",
      "votes": null
    },
    {
      "id": "454881",
      "postDate": "01/12/2019 12:29:06",
      "content": "<p>Validation is indeed a problem in this competition. For me, a good validation strategy is as follows:</p>\n\n<p>Let <strong>S</strong> be the matrix of probabilities such that <strong>S</strong>[i][j] is the probability that i-th image from  train and j-th image (also from train, but can be from test as well) are of the same whales. Now the trick is the understanding of a following idea: let N be the number of train images in total, which means that there are N^2 possible pairs of images. Per epoch, we process O(N) pairs, and therefore, assuming that there are even 1000 epochs, there are N^2 - 1000 * N pairs (which is a lot) that the model hasn't seen. Now, the good indication that the model can generalise well to unseen data is that it is more confident in its predictions pf this unseen data. This gives us a hint that you might want to watch for the std (standard deviation) of its predictions on the bounds (0 and 1). Thus, calculating 0.5-percentile and averaging the stds of <strong>S</strong> on both sides of this percentiles is a decent validation metric.</p>\n\n<p><strong>Note:</strong> in my experiments, this turned out not to be so good when the model if heavily overfitting, but this metric can easily be used to compare not-so-overfitting models without submitting to Kaggle.</p>\n\n<p><strong>Note 2:</strong> LB is still the best validation imho :)</p>",
      "rawMarkdown": "Validation is indeed a problem in this competition. For me, a good validation strategy is as follows:\n\nLet **S** be the matrix of probabilities such that **S**[i][j] is the probability that i-th image from  train and j-th image (also from train, but can be from test as well) are of the same whales. Now the trick is the understanding of a following idea: let N be the number of train images in total, which means that there are N^2 possible pairs of images. Per epoch, we process O(N) pairs, and therefore, assuming that there are even 1000 epochs, there are N^2 - 1000 * N pairs (which is a lot) that the model hasn't seen. Now, the good indication that the model can generalise well to unseen data is that it is more confident in its predictions pf this unseen data. This gives us a hint that you might want to watch for the std (standard deviation) of its predictions on the bounds (0 and 1). Thus, calculating 0.5-percentile and averaging the stds of **S** on both sides of this percentiles is a decent validation metric.\n\n**Note:** in my experiments, this turned out not to be so good when the model if heavily overfitting, but this metric can easily be used to compare not-so-overfitting models without submitting to Kaggle.\n\n**Note 2:** LB is still the best validation imho :)",
      "votes": null
    },
    {
      "id": "454932",
      "postDate": "01/12/2019 14:41:50",
      "content": "<p>Tanks for sharing. It is really useful!</p>",
      "rawMarkdown": "Tanks for sharing. It is really useful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454704,
      "author_name": "lanjunyelan",
      "author_url": "",
      "post_date": "01/12/2019 03:39:43",
      "content": "<p>I am also puzzled by this question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454881,
      "author_name": "vshakhray",
      "author_url": "",
      "post_date": "01/12/2019 12:29:06",
      "content": "<p>Validation is indeed a problem in this competition. For me, a good validation strategy is as follows:</p>\n\n<p>Let <strong>S</strong> be the matrix of probabilities such that <strong>S</strong>[i][j] is the probability that i-th image from  train and j-th image (also from train, but can be from test as well) are of the same whales. Now the trick is the understanding of a following idea: let N be the number of train images in total, which means that there are N^2 possible pairs of images. Per epoch, we process O(N) pairs, and therefore, assuming that there are even 1000 epochs, there are N^2 - 1000 * N pairs (which is a lot) that the model hasn't seen. Now, the good indication that the model can generalise well to unseen data is that it is more confident in its predictions pf this unseen data. This gives us a hint that you might want to watch for the std (standard deviation) of its predictions on the bounds (0 and 1). Thus, calculating 0.5-percentile and averaging the stds of <strong>S</strong> on both sides of this percentiles is a decent validation metric.</p>\n\n<p><strong>Note:</strong> in my experiments, this turned out not to be so good when the model if heavily overfitting, but this metric can easily be used to compare not-so-overfitting models without submitting to Kaggle.</p>\n\n<p><strong>Note 2:</strong> LB is still the best validation imho :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 454932,
          "author_name": "zhangjiescut",
          "author_url": "",
          "post_date": "01/12/2019 14:41:50",
          "content": "<p>Tanks for sharing. It is really useful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454090": "How to create validation data for simaese network? Is there any work related?",
    "454704": "I am also puzzled by this question.",
    "454881": "Validation is indeed a problem in this competition. For me, a good validation strategy is as follows:\n\nLet **S** be the matrix of probabilities such that **S**[i][j] is the probability that i-th image from  train and j-th image (also from train, but can be from test as well) are of the same whales. Now the trick is the understanding of a following idea: let N be the number of train images in total, which means that there are N^2 possible pairs of images. Per epoch, we process O(N) pairs, and therefore, assuming that there are even 1000 epochs, there are N^2 - 1000 * N pairs (which is a lot) that the model hasn't seen. Now, the good indication that the model can generalise well to unseen data is that it is more confident in its predictions pf this unseen data. This gives us a hint that you might want to watch for the std (standard deviation) of its predictions on the bounds (0 and 1). Thus, calculating 0.5-percentile and averaging the stds of **S** on both sides of this percentiles is a decent validation metric.\n\n**Note:** in my experiments, this turned out not to be so good when the model if heavily overfitting, but this metric can easily be used to compare not-so-overfitting models without submitting to Kaggle.\n\n**Note 2:** LB is still the best validation imho :)",
    "454932": "Tanks for sharing. It is really useful!"
  },
  "source": "meta"
}