{
  "id": 57755,
  "title": "All submissions using DBscan have exact same score (??)",
  "url": "/competitions/trackml-particle-identification/discussion/57755",
  "author_name": "",
  "post_date": "2018-05-28T16:08:28.704316400Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have made made 5 submissions using the basic DBscan kernel and I keep getting the same DBscan baseline score of .2078. I even tried using an event from train 2 instead of train 1 and the results are the same.  Each time I upload a submission the scoring feature hangs up with the message about getting a cup of coffee while extra time is spent on scoring and then coming back and refreshing the screen (F5). After I do the refresh  the scoring always completes and  I get another .2078 score.</p>\n\n<p>Can someone explain why the score is always the same even when different datasets are used  for training?</p>\n\n<p>tx</p>\n\n<p>Kickback</p>",
  "messages": [
    {
      "id": "334867",
      "postDate": "05/28/2018 16:08:28",
      "content": "<p>I have made made 5 submissions using the basic DBscan kernel and I keep getting the same DBscan baseline score of .2078. I even tried using an event from train 2 instead of train 1 and the results are the same.  Each time I upload a submission the scoring feature hangs up with the message about getting a cup of coffee while extra time is spent on scoring and then coming back and refreshing the screen (F5). After I do the refresh  the scoring always completes and  I get another .2078 score.</p>\n\n<p>Can someone explain why the score is always the same even when different datasets are used  for training?</p>\n\n<p>tx</p>\n\n<p>Kickback</p>",
      "rawMarkdown": "I have made made 5 submissions using the basic DBscan kernel and I keep getting the same DBscan baseline score of .2078. I even tried using an event from train 2 instead of train 1 and the results are the same.  Each time I upload a submission the scoring feature hangs up with the message about getting a cup of coffee while extra time is spent on scoring and then coming back and refreshing the screen (F5). After I do the refresh  the scoring always completes and  I get another .2078 score.\n\n   Can someone explain why the score is always the same even when different datasets are used  for training?\n\ntx\n\nKickback",
      "votes": null
    },
    {
      "id": "335386",
      "postDate": "05/29/2018 17:08:04",
      "content": "<p>\"This is a simple loop over the one-event actions: because of the use of DBScan, there is no actual training.\" DBSCAN is a deterministic algorithm; it doesn't \"learn\" anything, it simply makes predictions. So each time you run the data on the test set, you should expect the same output, regardless of what data you ran it with before.</p>\n\n<p>Edit: see <a href=\"http://scikit-learn.org/stable/modules/clustering.html#dbscan\">http://scikit-learn.org/stable/modules/clustering.html#dbscan</a> for more info</p>",
      "rawMarkdown": "\"This is a simple loop over the one-event actions: because of the use of DBScan, there is no actual training.\" DBSCAN is a deterministic algorithm; it doesn't \"learn\" anything, it simply makes predictions. So each time you run the data on the test set, you should expect the same output, regardless of what data you ran it with before.\n\nEdit: see http://scikit-learn.org/stable/modules/clustering.html#dbscan for more info",
      "votes": null
    },
    {
      "id": "335465",
      "postDate": "05/29/2018 19:26:39",
      "content": "<p>Niall,</p>\n\n<p>Thanks for your comments. I went back and read the comments in the kernel I used as a starting point (<a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark</a>) and my understanding now is that the \"Clusterer\" object processes data (either from a training dataset or the test dataset) so that clusters can be \"identified\"  using DBScan  and assigned labels. Clusterer does not get trained in any way, you just get to see the scores of your training data when it is used as the input for Clusterer (or you get to see the score of the test data after it has been used as input for Clusterer).  It is up to the user to improve  Clusterer to get a higher score if you intend to use the DBScan approach to identify clusters and give them labels.</p>\n\n<p>So my scores from Kaggle are all the same because Clusterer is always the same and the test data input to it is not changing.</p>\n\n<p>Am I stating this correctly? </p>\n\n<p>Thanks</p>\n\n<p>Kickback</p>",
      "rawMarkdown": "Niall,\n\nThanks for your comments. I went back and read the comments in the kernel I used as a starting point (https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark) and my understanding now is that the \"Clusterer\" object processes data (either from a training dataset or the test dataset) so that clusters can be \"identified\"  using DBScan  and assigned labels. Clusterer does not get trained in any way, you just get to see the scores of your training data when it is used as the input for Clusterer (or you get to see the score of the test data after it has been used as input for Clusterer).  It is up to the user to improve  Clusterer to get a higher score if you intend to use the DBScan approach to identify clusters and give them labels.\n\n  So my scores from Kaggle are all the same because Clusterer is always the same and the test data input to it is not changing.\n\nAm I stating this correctly? \n\nThanks\n\nKickback",
      "votes": null
    },
    {
      "id": "335496",
      "postDate": "05/29/2018 20:27:41",
      "content": "<p>Yes, that's correct. So you'll need to change the algorithm or modify the data to get a better score.</p>",
      "rawMarkdown": "Yes, that's correct. So you'll need to change the algorithm or modify the data to get a better score.",
      "votes": null
    },
    {
      "id": "335563",
      "postDate": "05/29/2018 22:47:31",
      "content": "<p>thanks again</p>",
      "rawMarkdown": "thanks again",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 335386,
      "author_name": "daemons",
      "author_url": "",
      "post_date": "05/29/2018 17:08:04",
      "content": "<p>\"This is a simple loop over the one-event actions: because of the use of DBScan, there is no actual training.\" DBSCAN is a deterministic algorithm; it doesn't \"learn\" anything, it simply makes predictions. So each time you run the data on the test set, you should expect the same output, regardless of what data you ran it with before.</p>\n\n<p>Edit: see <a href=\"http://scikit-learn.org/stable/modules/clustering.html#dbscan\">http://scikit-learn.org/stable/modules/clustering.html#dbscan</a> for more info</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 335465,
      "author_name": "bkosar1640",
      "author_url": "",
      "post_date": "05/29/2018 19:26:39",
      "content": "<p>Niall,</p>\n\n<p>Thanks for your comments. I went back and read the comments in the kernel I used as a starting point (<a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark</a>) and my understanding now is that the \"Clusterer\" object processes data (either from a training dataset or the test dataset) so that clusters can be \"identified\"  using DBScan  and assigned labels. Clusterer does not get trained in any way, you just get to see the scores of your training data when it is used as the input for Clusterer (or you get to see the score of the test data after it has been used as input for Clusterer).  It is up to the user to improve  Clusterer to get a higher score if you intend to use the DBScan approach to identify clusters and give them labels.</p>\n\n<p>So my scores from Kaggle are all the same because Clusterer is always the same and the test data input to it is not changing.</p>\n\n<p>Am I stating this correctly? </p>\n\n<p>Thanks</p>\n\n<p>Kickback</p>",
      "votes": null,
      "replies": [
        {
          "id": 335496,
          "author_name": "daemons",
          "author_url": "",
          "post_date": "05/29/2018 20:27:41",
          "content": "<p>Yes, that's correct. So you'll need to change the algorithm or modify the data to get a better score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 335563,
      "author_name": "bkosar1640",
      "author_url": "",
      "post_date": "05/29/2018 22:47:31",
      "content": "<p>thanks again</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "334867": "I have made made 5 submissions using the basic DBscan kernel and I keep getting the same DBscan baseline score of .2078. I even tried using an event from train 2 instead of train 1 and the results are the same.  Each time I upload a submission the scoring feature hangs up with the message about getting a cup of coffee while extra time is spent on scoring and then coming back and refreshing the screen (F5). After I do the refresh  the scoring always completes and  I get another .2078 score.\n\n   Can someone explain why the score is always the same even when different datasets are used  for training?\n\ntx\n\nKickback",
    "335386": "\"This is a simple loop over the one-event actions: because of the use of DBScan, there is no actual training.\" DBSCAN is a deterministic algorithm; it doesn't \"learn\" anything, it simply makes predictions. So each time you run the data on the test set, you should expect the same output, regardless of what data you ran it with before.\n\nEdit: see http://scikit-learn.org/stable/modules/clustering.html#dbscan for more info",
    "335465": "Niall,\n\nThanks for your comments. I went back and read the comments in the kernel I used as a starting point (https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark) and my understanding now is that the \"Clusterer\" object processes data (either from a training dataset or the test dataset) so that clusters can be \"identified\"  using DBScan  and assigned labels. Clusterer does not get trained in any way, you just get to see the scores of your training data when it is used as the input for Clusterer (or you get to see the score of the test data after it has been used as input for Clusterer).  It is up to the user to improve  Clusterer to get a higher score if you intend to use the DBScan approach to identify clusters and give them labels.\n\n  So my scores from Kaggle are all the same because Clusterer is always the same and the test data input to it is not changing.\n\nAm I stating this correctly? \n\nThanks\n\nKickback",
    "335496": "Yes, that's correct. So you'll need to change the algorithm or modify the data to get a better score.",
    "335563": "thanks again"
  },
  "source": "meta"
}