{
  "id": 76651,
  "title": "Score in LB and CV with pytorch",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76651",
  "author_name": "",
  "post_date": "2019-01-05T09:48:14.030728Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In first time, I try to use keras to bulid model. It can get the 0.698 in LB. However, I find the keras is unstable to score. Thus, I move to use Pytorch. This is my first time to use Pytorch and find many initialization different with keras such LSTM. After resetting the initialization, the new kernel with Pytorch cannot get same score in LB and lower than keras. Is LB sensitive with random seed? Or I ignore some difference bewteen keras and Pytorch? Many kagglers say they trust CV rather than LB, so I think I will persist in CV. \nIn addition, I hope to find some approachs to improve my CV score. My best CV score is just 0.685. There are many kaggler can get 0.69 or better, so I need some suggestion to improve my score.\nMy main approach:\n1. cleaning text: convert misspelling words, symbols and convert number to #.\n2. statistics features\n3. 5 folds\n3. model: two layers with bi-rnn + attention (I try to use capsule layer, but it takes too much times and cannot improve my CV scores.)</p>\n\n<p>Thank you for watching and your suggestions.</p>",
  "messages": [
    {
      "id": "450576",
      "postDate": "01/05/2019 09:48:14",
      "content": "<p>In first time, I try to use keras to bulid model. It can get the 0.698 in LB. However, I find the keras is unstable to score. Thus, I move to use Pytorch. This is my first time to use Pytorch and find many initialization different with keras such LSTM. After resetting the initialization, the new kernel with Pytorch cannot get same score in LB and lower than keras. Is LB sensitive with random seed? Or I ignore some difference bewteen keras and Pytorch? Many kagglers say they trust CV rather than LB, so I think I will persist in CV. \nIn addition, I hope to find some approachs to improve my CV score. My best CV score is just 0.685. There are many kaggler can get 0.69 or better, so I need some suggestion to improve my score.\nMy main approach:\n1. cleaning text: convert misspelling words, symbols and convert number to #.\n2. statistics features\n3. 5 folds\n3. model: two layers with bi-rnn + attention (I try to use capsule layer, but it takes too much times and cannot improve my CV scores.)</p>\n\n<p>Thank you for watching and your suggestions.</p>",
      "rawMarkdown": "In first time, I try to use keras to bulid model. It can get the 0.698 in LB. However, I find the keras is unstable to score. Thus, I move to use Pytorch. This is my first time to use Pytorch and find many initialization different with keras such LSTM. After resetting the initialization, the new kernel with Pytorch cannot get same score in LB and lower than keras. Is LB sensitive with random seed? Or I ignore some difference bewteen keras and Pytorch? Many kagglers say they trust CV rather than LB, so I think I will persist in CV. \nIn addition, I hope to find some approachs to improve my CV score. My best CV score is just 0.685. There are many kaggler can get 0.69 or better, so I need some suggestion to improve my score.\nMy main approach:\n1. cleaning text: convert misspelling words, symbols and convert number to #.\n2. statistics features\n3. 5 folds\n3. model: two layers with bi-rnn + attention (I try to use capsule layer, but it takes too much times and cannot improve my CV scores.)\n\nThank you for watching and your suggestions.",
      "votes": null
    },
    {
      "id": "450631",
      "postDate": "01/05/2019 12:34:21",
      "content": "<p>Deeper and bigger model structure can improve CV score , but only CV......</p>",
      "rawMarkdown": "Deeper and bigger model structure can improve CV score , but only CV......",
      "votes": null
    },
    {
      "id": "451007",
      "postDate": "01/06/2019 08:52:41",
      "content": "<p>I think some pre-process could improve the score too. However, I cannot find any other approaches for data cleaning. I need some ideas.</p>",
      "rawMarkdown": "I think some pre-process could improve the score too. However, I cannot find any other approaches for data cleaning. I need some ideas.",
      "votes": null
    },
    {
      "id": "451303",
      "postDate": "01/06/2019 20:00:07",
      "content": "<p>From what I have seen Preprocessing hasn't helped much with the scores. I would rather work on different models or statistical features. </p>",
      "rawMarkdown": "From what I have seen Preprocessing hasn't helped much with the scores. I would rather work on different models or statistical features.",
      "votes": null
    },
    {
      "id": "452698",
      "postDate": "01/09/2019 04:01:38",
      "content": "<p>What's your local and LB score with Pytorch instead of keras? Is there a big difference? </p>",
      "rawMarkdown": "What's your local and LB score with Pytorch instead of keras? Is there a big difference?",
      "votes": null
    },
    {
      "id": "452715",
      "postDate": "01/09/2019 04:30:28",
      "content": "<p>The best local CV is 0.683 and LB is 0.698. I guess keras randomness luck for me, but I want to use Pytorch to finish the competition beacuse it saves times to do more things such as data preprocess, \nfeature engineering. They can improve my score and make me confident for my score.</p>",
      "rawMarkdown": "The best local CV is 0.683 and LB is 0.698. I guess keras randomness luck for me, but I want to use Pytorch to finish the competition beacuse it saves times to do more things such as data preprocess, \nfeature engineering. They can improve my score and make me confident for my score.",
      "votes": null
    },
    {
      "id": "457424",
      "postDate": "01/17/2019 11:31:44",
      "content": "<p>How to use statistical features?  Can you take a example for me?\nI just know clean text and use word embedding into the model.\n Thanks. </p>",
      "rawMarkdown": "How to use statistical features?  Can you take a example for me?\nI just know clean text and use word embedding into the model.\n Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 450631,
      "author_name": "noxuslol",
      "author_url": "",
      "post_date": "01/05/2019 12:34:21",
      "content": "<p>Deeper and bigger model structure can improve CV score , but only CV......</p>",
      "votes": null,
      "replies": [
        {
          "id": 451007,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "01/06/2019 08:52:41",
          "content": "<p>I think some pre-process could improve the score too. However, I cannot find any other approaches for data cleaning. I need some ideas.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 451303,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "01/06/2019 20:00:07",
          "content": "<p>From what I have seen Preprocessing hasn't helped much with the scores. I would rather work on different models or statistical features. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452698,
      "author_name": "wxytalent",
      "author_url": "",
      "post_date": "01/09/2019 04:01:38",
      "content": "<p>What's your local and LB score with Pytorch instead of keras? Is there a big difference? </p>",
      "votes": null,
      "replies": [
        {
          "id": 452715,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "01/09/2019 04:30:28",
          "content": "<p>The best local CV is 0.683 and LB is 0.698. I guess keras randomness luck for me, but I want to use Pytorch to finish the competition beacuse it saves times to do more things such as data preprocess, \nfeature engineering. They can improve my score and make me confident for my score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 457424,
      "author_name": "karwik",
      "author_url": "",
      "post_date": "01/17/2019 11:31:44",
      "content": "<p>How to use statistical features?  Can you take a example for me?\nI just know clean text and use word embedding into the model.\n Thanks. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "450576": "In first time, I try to use keras to bulid model. It can get the 0.698 in LB. However, I find the keras is unstable to score. Thus, I move to use Pytorch. This is my first time to use Pytorch and find many initialization different with keras such LSTM. After resetting the initialization, the new kernel with Pytorch cannot get same score in LB and lower than keras. Is LB sensitive with random seed? Or I ignore some difference bewteen keras and Pytorch? Many kagglers say they trust CV rather than LB, so I think I will persist in CV. \nIn addition, I hope to find some approachs to improve my CV score. My best CV score is just 0.685. There are many kaggler can get 0.69 or better, so I need some suggestion to improve my score.\nMy main approach:\n1. cleaning text: convert misspelling words, symbols and convert number to #.\n2. statistics features\n3. 5 folds\n3. model: two layers with bi-rnn + attention (I try to use capsule layer, but it takes too much times and cannot improve my CV scores.)\n\nThank you for watching and your suggestions.",
    "450631": "Deeper and bigger model structure can improve CV score , but only CV......",
    "451007": "I think some pre-process could improve the score too. However, I cannot find any other approaches for data cleaning. I need some ideas.",
    "451303": "From what I have seen Preprocessing hasn't helped much with the scores. I would rather work on different models or statistical features.",
    "452698": "What's your local and LB score with Pytorch instead of keras? Is there a big difference?",
    "452715": "The best local CV is 0.683 and LB is 0.698. I guess keras randomness luck for me, but I want to use Pytorch to finish the competition beacuse it saves times to do more things such as data preprocess, \nfeature engineering. They can improve my score and make me confident for my score.",
    "457424": "How to use statistical features?  Can you take a example for me?\nI just know clean text and use word embedding into the model.\n Thanks."
  },
  "source": "meta"
}