{"id":33410,"date":"2021-08-10T13:58:12","date_gmt":"2021-08-10T17:58:12","guid":{"rendered":"https:\/\/www.bu.edu\/cise\/?page_id=33410"},"modified":"2022-01-23T23:48:53","modified_gmt":"2022-01-24T04:48:53","slug":"cise-seminar-april-1-2019-ashok-cutkosky-google","status":"publish","type":"page","link":"https:\/\/www.bu.edu\/cise\/cise-seminars-archive\/spring-2019-2\/cise-seminar-april-1-2019-ashok-cutkosky-google\/","title":{"rendered":"CISE Seminar: April 1, 2019 &#8211; Ashok Cutkosky, Google"},"content":{"rendered":"<p>BU Photonics Building<br \/>\n8 St. Mary&#8217;s Street, <span style=\"color: #ff9900;\">PHO 339<\/span><br \/>\n<span style=\"color: #ff9900;\">1:30pm-2:30pm<\/span><\/p>\n<h2><span style=\"color: #000000;\"><strong><img loading=\"lazy\" src=\"\/cise\/files\/2019\/03\/ashokcutkosky_pic1-e1553609443134-547x636.jpg\" alt=\"\" width=\"150\" height=\"175\" class=\"alignleft wp-image-24851\" srcset=\"https:\/\/www.bu.edu\/cise\/files\/2019\/03\/ashokcutkosky_pic1-e1553609443134-547x636.jpg 547w, https:\/\/www.bu.edu\/cise\/files\/2019\/03\/ashokcutkosky_pic1-e1553609443134.jpg 600w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/>Ashok Cutkosky<br \/>\n<\/strong><\/span><span style=\"color: #000000;\">Google<\/span><\/h2>\n<h3><span style=\"color: #000000;\">Preconditioned Online Learning without Preconditioning<\/span><\/h3>\n<p><span>In stochastic convex optimization, the rate of convergence is often dominated by the variance in the gradients. In order to combat this, one can make use of the covariance matrix of the gradients. This information allows an optimization algorithm to ignore irrelevant &#8220;noise directions&#8221;, which can lead to a much lower &#8220;effective variance&#8221; (in the best case reducing from the trace of the covariance matrix to its smallest eigenvalue). In an apparent effort to promote confusion with Newton&#8217;s method for non-stochastic optimization, this is often called &#8220;preconditioning&#8221;. <\/span><\/p>\n<p><span>Unfortunately (similar to Newton&#8217;s method), preconditioning requires computing, storing, and inverting a matrix, which takes time and space at best quadratic in the problem dimension rather than the linear time of gradient descent. This is prohibitive for modern high-dimensional problems, so practitioners typically resort to linear time algorithms such as AdaGrad which approximate the covariance matrix by just its diagonal entries. In this talk, I will describe a technique that attempts to achieve full preconditioning while using only linear time and space. This technique never does worse than gradient descent or diagonal techniques such as AdaGrad, and in certain settings, it actually surpasses the guarantees provided by preconditioning algorithms that use the entire covariance matrix.<\/span><\/p>\n<p><span><b>Ashok Cutkosky\u00a0<\/b>is<span style=\"color: #333333;\">\u00a0currently a research scientist at Google. He obtained a Ph.D. in computer science from Stanford in 2018, a Masters in Medicine from Stanford in 2016, and a bachelors in Mathematics from Harvard in 2013. He is interested in optimization theory, and in adaptive learning algorithms in particular. He received the best student paper award at COLT in 2017, and he also enjoys card magic.<\/span><\/span><\/p>\n<p><span style=\"color: #000000;\">Faculty Host: Yannis Paschalidis<\/span><br \/>\n<span style=\"color: #000000;\">Student Hosts:\u00a0Zhenxun Zhuang and Artin Spiridonoff<\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>BU Photonics Building 8 St. Mary&#8217;s Street, PHO 339 1:30pm-2:30pm Ashok Cutkosky Google Preconditioned Online Learning without Preconditioning In stochastic convex optimization, the rate of convergence is often dominated by the variance in the gradients. In order to combat this, one can make use of the covariance matrix of the gradients. This information allows an [&hellip;]<\/p>\n","protected":false},"author":18553,"featured_media":0,"parent":35320,"menu_order":5,"comment_status":"closed","ping_status":"closed","template":"","meta":[],"_links":{"self":[{"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/pages\/33410"}],"collection":[{"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/users\/18553"}],"replies":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/comments?post=33410"}],"version-history":[{"count":1,"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/pages\/33410\/revisions"}],"predecessor-version":[{"id":33411,"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/pages\/33410\/revisions\/33411"}],"up":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/pages\/35320"}],"wp:attachment":[{"href":"https:\/\/www.bu.edu\/cise\/wp-json\/wp\/v2\/media?parent=33410"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}