{
  "id": 54919,
  "title": "LGBM Tuning Question - Min_child_samples",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54919",
  "author_name": "",
  "post_date": "2018-04-19T17:25:27.725397800Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>What is the difference between min_child_samples and min_child_weight </p>\n\n<p>min_child_samples , Is it the minimum number of leaves required in a leaf to process the tree further? If so is the idea that by keeping it at a larger value we can avoid overfitting? </p>\n\n<p>Kind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100. </p>\n\n<p>Where does min_child_weight fit into all this?</p>\n\n<p>Regards\nShanth</p>",
  "messages": [
    {
      "id": "316696",
      "postDate": "04/19/2018 17:25:27",
      "content": "<p>What is the difference between min_child_samples and min_child_weight </p>\n\n<p>min_child_samples , Is it the minimum number of leaves required in a leaf to process the tree further? If so is the idea that by keeping it at a larger value we can avoid overfitting? </p>\n\n<p>Kind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100. </p>\n\n<p>Where does min_child_weight fit into all this?</p>\n\n<p>Regards\nShanth</p>",
      "rawMarkdown": "What is the difference between min_child_samples and min_child_weight \n\nmin_child_samples , Is it the minimum number of leaves required in a leaf to process the tree further? If so is the idea that by keeping it at a larger value we can avoid overfitting? \n\nKind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100. \n\nWhere does min_child_weight fit into all this?\n\nRegards\nShanth",
      "votes": null
    },
    {
      "id": "316908",
      "postDate": "04/20/2018 07:08:31",
      "content": "<p>Suspect it be something about the sample weight.</p>",
      "rawMarkdown": "Suspect it be something about the sample weight.",
      "votes": null
    },
    {
      "id": "316924",
      "postDate": "04/20/2018 07:38:30",
      "content": "<p>Thanks for response Fei.  Could you elaborate on how it relates to overfit? Do point me to some online resources if you have any.</p>",
      "rawMarkdown": "Thanks for response Fei.  Could you elaborate on how it relates to overfit? Do point me to some online resources if you have any.",
      "votes": null
    },
    {
      "id": "316928",
      "postDate": "04/20/2018 07:42:56",
      "content": "<p>Sorry that was just from my hunch. Think you need the link in the comment of <a href=\"/pranav\">@pranav</a> of <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696#310126\">this post</a></p>",
      "rawMarkdown": "Sorry that was just from my hunch. Think you need the link in the comment of @pranav of [this post](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696#310126)",
      "votes": null
    },
    {
      "id": "316996",
      "postDate": "04/20/2018 12:20:55",
      "content": "<p>min_child_samples is minimal amount of data needed in a leaf</p>\n\n<blockquote>\n  <p>Kind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100.</p>\n</blockquote>\n\n<p>Not sure about this (i.e. algorithm actually stopping), my interpretation is that algorithm will not create a leaf if there is less than 100 samples.</p>",
      "rawMarkdown": "min_child_samples is minimal amount of data needed in a leaf\n\n&gt; Kind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100.\n\nNot sure about this (i.e. algorithm actually stopping), my interpretation is that algorithm will not create a leaf if there is less than 100 samples.",
      "votes": null
    },
    {
      "id": "317000",
      "postDate": "04/20/2018 12:30:15",
      "content": "<p>@ Konchar - Thanks for the explanation</p>",
      "rawMarkdown": "Konchar - Thanks for the explanation",
      "votes": null
    },
    {
      "id": "317001",
      "postDate": "04/20/2018 12:30:31",
      "content": "<p>@ Thanks Fei</p>",
      "rawMarkdown": "Thanks Fei",
      "votes": null
    },
    {
      "id": "317004",
      "postDate": "04/20/2018 12:38:48",
      "content": "<p>You are welcome, and thank you for asking about min_child_weight because you actually made me explore the subject (I initially thought you were talking about scale_pos_weight). I'll let you know if I find anything useful about min_child_weight!</p>",
      "rawMarkdown": "You are welcome, and thank you for asking about min_child_weight because you actually made me explore the subject (I initially thought you were talking about scale_pos_weight). I'll let you know if I find anything useful about min_child_weight!",
      "votes": null
    },
    {
      "id": "317014",
      "postDate": "04/20/2018 13:15:17",
      "content": "<p>@ Konchar.   Sure thanks :)</p>",
      "rawMarkdown": "Konchar.   Sure thanks :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 316908,
      "author_name": "enfeizhan",
      "author_url": "",
      "post_date": "04/20/2018 07:08:31",
      "content": "<p>Suspect it be something about the sample weight.</p>",
      "votes": null,
      "replies": [
        {
          "id": 316924,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "04/20/2018 07:38:30",
          "content": "<p>Thanks for response Fei.  Could you elaborate on how it relates to overfit? Do point me to some online resources if you have any.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316928,
          "author_name": "enfeizhan",
          "author_url": "",
          "post_date": "04/20/2018 07:42:56",
          "content": "<p>Sorry that was just from my hunch. Think you need the link in the comment of <a href=\"/pranav\">@pranav</a> of <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696#310126\">this post</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 317001,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "04/20/2018 12:30:31",
          "content": "<p>@ Thanks Fei</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 316996,
      "author_name": "konchar",
      "author_url": "",
      "post_date": "04/20/2018 12:20:55",
      "content": "<p>min_child_samples is minimal amount of data needed in a leaf</p>\n\n<blockquote>\n  <p>Kind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100.</p>\n</blockquote>\n\n<p>Not sure about this (i.e. algorithm actually stopping), my interpretation is that algorithm will not create a leaf if there is less than 100 samples.</p>",
      "votes": null,
      "replies": [
        {
          "id": 317000,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "04/20/2018 12:30:15",
          "content": "<p>@ Konchar - Thanks for the explanation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 317004,
          "author_name": "konchar",
          "author_url": "",
          "post_date": "04/20/2018 12:38:48",
          "content": "<p>You are welcome, and thank you for asking about min_child_weight because you actually made me explore the subject (I initially thought you were talking about scale_pos_weight). I'll let you know if I find anything useful about min_child_weight!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 317014,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "04/20/2018 13:15:17",
          "content": "<p>@ Konchar.   Sure thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "316696": "What is the difference between min_child_samples and min_child_weight \n\nmin_child_samples , Is it the minimum number of leaves required in a leaf to process the tree further? If so is the idea that by keeping it at a larger value we can avoid overfitting? \n\nKind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100. \n\nWhere does min_child_weight fit into all this?\n\nRegards\nShanth",
    "316908": "Suspect it be something about the sample weight.",
    "316924": "Thanks for response Fei.  Could you elaborate on how it relates to overfit? Do point me to some online resources if you have any.",
    "316928": "Sorry that was just from my hunch. Think you need the link in the comment of @pranav of [this post](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696#310126)",
    "316996": "min_child_samples is minimal amount of data needed in a leaf\n\n&gt; Kind of like saying the algorithm will stop when a leaf has 95 samples for a min_child_samples = 100.\n\nNot sure about this (i.e. algorithm actually stopping), my interpretation is that algorithm will not create a leaf if there is less than 100 samples.",
    "317000": "Konchar - Thanks for the explanation",
    "317001": "Thanks Fei",
    "317004": "You are welcome, and thank you for asking about min_child_weight because you actually made me explore the subject (I initially thought you were talking about scale_pos_weight). I'll let you know if I find anything useful about min_child_weight!",
    "317014": "Konchar.   Sure thanks :)"
  },
  "source": "meta"
}