{"id":22799,"date":"2021-03-10T14:18:15","date_gmt":"2021-03-10T19:18:15","guid":{"rendered":"https:\/\/www.bu.edu\/hic\/?p=22799"},"modified":"2025-12-09T16:30:56","modified_gmt":"2025-12-09T21:30:56","slug":"machine-learning-models-have-to-memorize-sensitive-information","status":"publish","type":"post","link":"https:\/\/www.bu.edu\/hic\/2021\/03\/10\/machine-learning-models-have-to-memorize-sensitive-information\/","title":{"rendered":"Machine learning models have to memorize sensitive information"},"content":{"rendered":"<p>BY: GINA MANTICA<\/p>\n<p>You might\u2019ve encountered a machine learning model while typing an email or text that starts to automatically fill in the rest of a word, phrase, or sentence. Machine learning models for sentence prediction can not only complete sentences, but they can also analyze sentence structure and predict grammar. However, Hariri Institute Research Fellow Adam Smith, along with BU researchers Gavin Brown and Mark Bun and colleagues at Apple, found that models for problems like this cannot function without memorizing data that is oftentimes sensitive.<\/p>\n<p><a href=\"https:\/\/arxiv.org\/abs\/2012.06421\">Their paper<\/a> was recently accepted by the <a href=\"http:\/\/acm-stoc.org\/\">2021 Association for Computing Machinery (ACM) Symposium on Theory of Computing<\/a>.<\/p>\n<p>Machine learning models extract relevant information from examples, or training set data, and produce a prediction algorithm based on that information. But the most accurate models sometimes memorize what seems like irrelevant information, even at the expense of an individual\u2019s privacy. Smith\u2019s findings suggest that the model\u2019s performance and an individual\u2019s right to privacy can be fundamentally at odds.<\/p>\n<figure id=\"attachment_22804\" aria-describedby=\"caption-attachment-22804\" style=\"width: 410px\" class=\"wp-caption alignleft\"><img loading=\"lazy\" src=\"\/hic\/files\/2021\/03\/Smith-Adam-636x636.jpg\" alt=\"\" width=\"400\" height=\"400\" class=\"wp-image-22804\" srcset=\"https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-636x636.jpg 636w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-1024x1024.jpg 1024w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-150x150.jpg 150w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-768x768.jpg 768w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-1536x1536.jpg 1536w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-2048x2048.jpg 2048w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-700x700.jpg 700w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-189x189.jpg 189w, https:\/\/www.bu.edu\/hic\/files\/2021\/03\/Smith-Adam-100x100.jpg 100w\" sizes=\"(max-width: 400px) 100vw, 400px\" \/><figcaption id=\"caption-attachment-22804\" class=\"wp-caption-text\">Adam Smith, Hariri Institute Research Fellow and Professor in Computer Science, discovered that memorization is necessary for machines to learn.<\/figcaption><\/figure>\n<p>Through a detailed set of theoretical problems and mathematical proofs, Smith and colleagues find that any training algorithm with a limited-size data set must store the complete details of several data points to perform well. Smaller data sets are more likely to contain many outliers\u2014data points that are very different from the rest of the data set. Each such data point is rich in information, but it\u2019s hard for the model, given only one of them, to know which aspects of the data point are worth remembering and which are not. The model therefore must remember lots of extraneous details about the data in order to make accurate predictions.<\/p>\n<p>However, if there is a lot of training data, the machine learning model can achieve the same level of accuracy without memorizing as much individual information about data points. This is because larger training sets likely contain many examples of any given kind, allowing the model to hone in on and retain only essential features. Ironically, the more information the model is given, the less specific detail the model needs to retain.<\/p>\n<p>The researchers\u2019 work suggests that machine learning models have to memorize all the information about outliers, in particular, for models to make accurate predictions&#8211;getting rid of any information increases errors. \u201cWe showed that this phenomenon is unavoidable,\u201d said Smith, \u201cThere is no way of training machine learning models that gets around memorizing specific examples of the training set data.\u201d<\/p>\n<p>As machine learning models are implemented widely, their utility will need to be weighed carefully against their risk to privacy.<\/p>\n<p><em><\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>BY: GINA MANTICA You might\u2019ve encountered a machine learning model while typing an email or text that starts to automatically fill in the rest of a word, phrase, or sentence. Machine learning models for sentence prediction can not only complete sentences, but they can also analyze sentence structure and predict grammar. However, Hariri Institute Research [&hellip;]<\/p>\n","protected":false},"author":18687,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[11716],"tags":[],"_links":{"self":[{"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/posts\/22799"}],"collection":[{"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/users\/18687"}],"replies":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/comments?post=22799"}],"version-history":[{"count":8,"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/posts\/22799\/revisions"}],"predecessor-version":[{"id":41247,"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/posts\/22799\/revisions\/41247"}],"wp:attachment":[{"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/media?parent=22799"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/categories?post=22799"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.bu.edu\/hic\/wp-json\/wp\/v2\/tags?post=22799"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}