{
  "id": 64988,
  "title": "Notes from my submission",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/64988",
  "author_name": "macfarll",
  "post_date": "2018-09-05T02:07:23.864000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.</p>\n\n<p>Most of the time I spent on this competition was working with my neural network, either trying to get the network to converge, or trying to overcome the network's tendency to predict all ads won't be attributed due to how rare attributed values were in the dataset. </p>\n\n<p>My solution to overcome these issues was to create a script for subsampling the data, and to train the network with over-represented attributed values. A network trained on oversampled attributed values tended to produce too many attributed predicitions. To account for this, I only let the algorithm partially converge, then decreased the amount of attributed values in the training set. </p>\n\n<p>My strategy of iteratively decreasing the amount of attributed values worked, and produced a network that converged and predicted a reasonable amount of each value, but my score wasn't very impressive. </p>\n\n<p>My solution fell short of the publicly available solutions, but offered slight lift to them when both solutions were ensembled.</p>\n\n<p>After reading suggestions from other participants and seeing the better solutions, the obvious shortcoming to my solution was a lack of features, and too much of a focus on the design of the network and hyper-parameter tuning. </p>",
  "messages": [
    {
      "id": 381689,
      "postDate": "2018-09-05T02:07:23.863Z",
      "content": "<p>So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.</p>\n\n<p>Most of the time I spent on this competition was working with my neural network, either trying to get the network to converge, or trying to overcome the network's tendency to predict all ads won't be attributed due to how rare attributed values were in the dataset. </p>\n\n<p>My solution to overcome these issues was to create a script for subsampling the data, and to train the network with over-represented attributed values. A network trained on oversampled attributed values tended to produce too many attributed predicitions. To account for this, I only let the algorithm partially converge, then decreased the amount of attributed values in the training set. </p>\n\n<p>My strategy of iteratively decreasing the amount of attributed values worked, and produced a network that converged and predicted a reasonable amount of each value, but my score wasn't very impressive. </p>\n\n<p>My solution fell short of the publicly available solutions, but offered slight lift to them when both solutions were ensembled.</p>\n\n<p>After reading suggestions from other participants and seeing the better solutions, the obvious shortcoming to my solution was a lack of features, and too much of a focus on the design of the network and hyper-parameter tuning. </p>",
      "rawMarkdown": "So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.\n\nMost of the time I spent on this competition was working with my neural network, either trying to get the network to converge, or trying to overcome the network's tendency to predict all ads won't be attributed due to how rare attributed values were in the dataset. \n\nMy solution to overcome these issues was to create a script for subsampling the data, and to train the network with over-represented attributed values. A network trained on oversampled attributed values tended to produce too many attributed predicitions. To account for this, I only let the algorithm partially converge, then decreased the amount of attributed values in the training set. \n\nMy strategy of iteratively decreasing the amount of attributed values worked, and produced a network that converged and predicted a reasonable amount of each value, but my score wasn't very impressive. \n\nMy solution fell short of the publicly available solutions, but offered slight lift to them when both solutions were ensembled.\n\nAfter reading suggestions from other participants and seeing the better solutions, the obvious shortcoming to my solution was a lack of features, and too much of a focus on the design of the network and hyper-parameter tuning. ",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "381689": "So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.\n\nMost of the time I spent on this competition was working with my neural network, either trying to get the network to converge, or trying to overcome the network's tendency to predict all ads won't be attributed due to how rare attributed values were in the dataset. \n\nMy solution to overcome these issues was to create a script for subsampling the data, and to train the network with over-represented attributed values. A network trained on oversampled attributed values tended to produce too many attributed predicitions. To account for this, I only let the algorithm partially converge, then decreased the amount of attributed values in the training set. \n\nMy strategy of iteratively decreasing the amount of attributed values worked, and produced a network that converged and predicted a reasonable amount of each value, but my score wasn't very impressive. \n\nMy solution fell short of the publicly available solutions, but offered slight lift to them when both solutions were ensembled.\n\nAfter reading suggestions from other participants and seeing the better solutions, the obvious shortcoming to my solution was a lack of features, and too much of a focus on the design of the network and hyper-parameter tuning. "
  }
}